Information provision system, information provision method, and program
Patent Information
- Application Number
- PCT/JP2026/004332
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-21
- Filing Date
- 2026-02-06
- Publication Date
- 2026-08-27
Smart Images

Figure JP2026004332_27082026_PF_FP_ABST
Abstract
Description
Information providing system, information providing method, and program
[0001] The present disclosure generally relates to an information providing system, an information providing method, and a program. More specifically, the present disclosure relates to an information providing system, an information providing method, and a program that provide information about the situation of a person who conducts event activities such as business activities.
[0002] Patent Document 1 discloses a user cooperation system. In this user cooperation system, data obtained by detecting the real-time situation of a user from sensors worn by the user and a variety of sensors installed around each user is sent to a server. The context analysis unit of the processing unit of the server determines the degree of busyness of each user based on this data, and the cooperation control unit controls the communication between each user based on this degree of busyness.
[0003] Japanese Patent Application Laid-Open No. 2009-20672
[0004] By the way, the situation of a target person (for example, a subordinate in the workplace), who is a user who conducts event activities such as business activities, may be desired to be efficiently and favorably transmitted to other users (for example, a supervisor in the workplace). On the other hand, if the transmission is made by language, it may be an obstacle to the activity of the user who receives the transmission because the concentration is inhibited. Therefore, in order not to interfere with the activity of the user who receives the transmission, there may be a case where transmission by a natural perception and a provision that blends in with the place such as the workplace is desired so that the user can know by natural awareness.
[0005] The present disclosure has been made in view of the above reasons, and an object thereof is to provide an information providing system, an information providing method, and a program that can easily realize information providing that blends in with the place.
[0006] An information provision system according to one aspect of this disclosure provides information about the status of a target person, which is any user among a group of users. The information provision system comprises an information acquisition unit, an estimation unit, a sound source selection unit, and a sound output unit. The information acquisition unit acquires target person information, including activity information relating to at least one of the voices spoken by the target person, the actions taken by the target person, and the characters entered by the target person, during the target person's event activities. The estimation unit estimates at least one of the target person's emotions and level of activity based on the target person information. The sound source selection unit selects sound source data corresponding to the evaluation information based on the estimation result by the estimation unit from among a group of sound source data. The sound output unit provides the information by outputting sound based on the sound source data selected by the sound source selection unit from an acoustic device.
[0007] An information provision method according to one aspect of this disclosure is an information provision method executed by one or more processors that provides information about the status of a subject, which is any user among a plurality of users. The information provision method includes an information acquisition process, an estimation process, a sound source selection process, and a sound output process. In the information acquisition process, subject information is acquired, which includes activity information relating to at least one of the voices spoken by the subject, the actions taken by the subject, and the characters input by the subject during the subject's event activities. In the estimation process, at least one of the subject's emotions and activity level is estimated based on the subject information. In the sound source selection process, sound source data corresponding to the evaluation information based on the estimation result from the estimation process is selected from a plurality of sound source data. In the sound output process, the information is provided by outputting sound based on the sound source data selected in the sound source selection process from an audio device.
[0008] A program according to one aspect of this disclosure is a program that causes one or more processors to execute the information provision method.
[0009] Figure 1 is a block diagram of an information provision system according to an embodiment. Figure 2 is a conceptual diagram illustrating the relationship between multiple users (workers) who are members of an organization and the sound source data associated with each user in the above information provision system. Figure 3 is a conceptual diagram illustrating a predetermined space in the above information provision system where a target person (e.g., person D) performs event activities (work). Figure 4 is a conceptual diagram illustrating a predetermined space in the above information provision system where a user (e.g., person A) who receives information about the situation of a target person (e.g., person D) performs event activities (work). Figure 5 is a conceptual diagram of the single-event sound source storage area and the continuous sound source storage area within the sound source data storage unit in the above information provision system. Figure 6 is a sequence diagram illustrating operation example 1 in the above information provision system. Figure 7 is a sequence diagram illustrating operation example 2 in the above information provision system.
[0010] (Summary) The following describes the information provision system, information provision method, and program relating to embodiments and modifications, with reference to drawings. Note that the embodiments and modifications described below are only one of the various embodiments of this disclosure. Furthermore, the embodiments and modifications described below can be modified in various ways depending on the design, etc., as long as the objectives of this disclosure are achieved. It is also possible to combine the configurations of the modifications as appropriate. In addition, the figures described in this disclosure are schematic diagrams, and the ratios of the size and thickness of each component in the figures do not necessarily reflect the actual dimensional ratios.
[0011] An information provision system 100 according to one embodiment (see Figure 1) provides information about the status of a target person T1, which is any user 300 among a plurality of users 300. The status of target person T1 is estimated and evaluated based on information obtained during target person T1's event activities (target person information), and then output as a corresponding sound (information provision) as described below.
[0012] In the following example, it is assumed that the multiple users 300 are multiple workers belonging to a company, government office, or other organization, etc., and performing their duties at their workplace (mainly an office, but at home in the case of teleworking). In other words, it is assumed that "event activities" are work activities (labor activities), and that the circumstances of workers can be inferred. In this disclosure, "worker" refers to all persons who perform work. This may include not only those who receive compensation for their work, but also those who work without compensation.
[0013] However, User 300 is not limited to workers. User 300 may be, for example, a student, pupil, or child who studies at a school or cram school, in which case multiple Users 300 may include multiple learners, teachers, parents of learners, etc., and the "event activity" may be a learning activity. Alternatively, User 300 may be, for example, a person (player) who plays games, in which case multiple Users 300 may include multiple players, and the "event activity" may be a play activity.
[0014] As shown in Figure 1, the information provision system 100 comprises an information acquisition unit 1, an estimation unit 2, a sound source selection unit 3, and a sound output unit 4. The information acquisition unit 1 acquires subject information, including activity information relating to at least one of the voices spoken by subject T1, actions taken by subject T1, and characters input by subject T1, during subject T1's event activities (e.g., during work activities). The activity information may be information sensed by various sensors 6 such as a microphone 61, and / or information acquired from a cloud server 7 of a communication tool, although details will be described later. Examples of communication tools include email tools, chat tools, or online meeting application tools such as Teams® or Zoom®.
[0015] The estimation unit 2 estimates at least one of the subject T1's emotions and activity level based on the subject information. The sound source selection unit 3 selects sound source data corresponding to the evaluation information based on the estimation result by the estimation unit 2 from among multiple sound source data. The sound output unit 4 provides information by outputting sounds based on the sound source data selected by the sound source selection unit 3 from the sound device 5. The "sound" referred to here is, for example, a natural sound. Examples of natural sounds include bird songs, bird wing sounds, animal footsteps, animal calls (other than birds), insect sounds, stream sounds, wave sounds, rain sounds, waterfall sounds, wind sounds, leaf rustling, bonfire sounds, gusts of wind, and thunder. Multiple sound source data of such natural sounds are stored in advance in the storage unit S1 (sound source data storage unit S13).
[0016] According to the configuration of the information provision system 100, the situation of a target person T1 (e.g., a subordinate in the workplace), who is a user 300 performing event activities such as work activities, can be efficiently and effectively communicated to other users 300 (e.g., a supervisor in the workplace) as sound data (e.g., natural sounds). In short, efficient communication from a subordinate to a supervisor can be achieved through "sound." In particular, since the information is not provided through direct verbal communication from the target person T1 themselves, the likelihood of the receiving user 300 having their concentration disrupted and their activities hindered is lower compared to when the communication is done through direct verbal communication. As a result, other users 300 (e.g., supervisors in the workplace) can become aware of the target person T1's situation through natural awareness, and their actions / ideas are stimulated. Consequently, the information provision system 100 has the advantage of making it easier to provide information that is in harmony with the situation. In particular, if the 300 users are workers, providing information that is integrated into the work environment could lead to improvements in productivity, sales, customer satisfaction, and employee retention rates within the organization.
[0017] An information provision method according to one embodiment is an information provision method executed by one or more processors that provides information about the status of a subject T1, which is any user 300 among a plurality of users 300. The information provision method includes an information acquisition process, an estimation process, a sound source selection process, and a sound output process. In the information acquisition process, subject information is acquired, which includes activity information relating to at least one of the voices spoken by subject T1, the actions taken by subject T1, and the characters input by subject T1 during subject T1's event activities. In the estimation process, at least one of subject T1's emotions and activity level is estimated based on the subject information. In the sound source selection process, sound source data corresponding to the evaluation information based on the estimation results from the estimation process is selected from among a plurality of sound source data. In the sound output process, information is provided by outputting sound based on the sound source data selected in the sound source selection process from an audio device.
[0018] This method of providing information also has the advantage of making it easier to provide information that is in harmony with the context.
[0019] This information provision method is used on a computer system (information provision system 100). In other words, this information provision method can also be implemented as a computer program. A program according to one embodiment is a program that causes one or more processors to execute the above information provision method. The program may be recorded on a computer-readable non-temporary recording medium. Furthermore, a computer program product according to one embodiment includes a computer program that, when executed by one or more processors, realizes the processing (steps) of the above information provision method.
[0020] In the following, it is assumed that multiple functions of the information provision system 100 are implemented on a server. The server includes one or more server devices. If the server includes multiple server devices, a cloud (cloud computing) may be constructed using these multiple server devices.
[0021] Furthermore, the multiple functions of the information provision system 100 are not limited to being implemented entirely on the server. At least some of the functions of the information provision system 100 may be implemented on the edge side (for example, on an information terminal used by a user 300). For example, the functions of the information acquisition unit 1 may be implemented on the edge side.
[0022] Unless otherwise specified, the "information terminals" described below are assumed to be smartphones, tablet devices, laptop computers, or desktop computers used by user 300. It is assumed that these information terminals are provided by the company to user 300, who is an employee, for the purpose of performing their duties. The terminal ID information for each information terminal is associated with the personal information of user 300 (name, employee number, and email address, etc.) and managed in a database (user information storage unit S14).
[0023] (Embodiment) (1) Overall Configuration The overall configuration of the labor management system, including the information provision system 100 and its peripheral configuration according to the embodiment, will be described below with reference to Figures 1 to 5.
[0024] As shown in Figure 1, the labor management system comprises an information provision system 100, one or more acoustic devices 5 (one in the illustrated example), one or more sensors 6 (four in the illustrated example), and a communication tool (cloud server 7).
[0025] The information provision system 100 provides information about the status of target person T1, which is any user 300 from among multiple users 300 (see Figure 2: workers). Here, we will explain using four workers (users 301-304) within a single organization (an organization is a department or division, for example, the Human Resources Department) within a company as an example. User 301 is Mr. A (position: department head) in the Human Resources Department, user 302 is Mr. B (position: section chief) in the Human Resources Department, user 303 is Mr. C (regular employee, senior to Mr. D) in the Human Resources Department, and user 304 is Mr. D (regular employee, newcomer) in the Human Resources Department. In the following explanation, target person T1 may be set as user 304 (newcomer Mr. D), but each of users 301-304 can be target person T1. Note that the organization is not limited to a department or division, but may also be a project team formed for a limited period of time.
[0026] As shown in Figure 1, the information provision system 100 comprises an information acquisition unit 1, an estimation unit 2, a sound source selection unit 3, a sound output unit 4, an evaluation unit H1, and a storage unit S1. The information provision system 100 is implemented, for example, by a computer system having one or more processors and one or more memories. In other words, multiple functions of the information provision system 100 (functions of the information acquisition unit 1, estimation unit 2, sound source selection unit 3, sound output unit 4, evaluation unit H1, etc.) are realized by one or more processors executing a program recorded in memory. The program may be pre-recorded in memory, provided via a telecommunication line such as the Internet, or provided on a non-temporary recording medium such as a memory card.
[0027] As described above, the various functions of the information provision system 100 are implemented, for example, on a server. The information provision system 100 can communicate with various sensors 6, a cloud server 7, and an acoustic device 5 via a wide-area communication network such as the Internet.
[0028] Detailed explanations of each function in the information provision system 100 will be provided later.
[0029] One or more sensors 6 are an example of means for obtaining "person information" that the information provision system 100 uses to estimate / evaluate the situation of the person T1. One or more sensors 6 are installed / used within a predetermined space 200 (see Figures 3 and 4) where the user 300 (person T1) conducts event activities (in this case, work activities). In other words, the sensors 6 may be installed within the predetermined space 200, or the user 300 may wear them on their body and use them within the predetermined space 200. One or more sensors 6 preferably include at least one of the four sensors 6: a microphone 61, a motion sensor 62, a position sensor 63, and a speech sensor 64, and here, as an example, all four are included.
[0030] In Figure 3, the designated space 200 is the workspace where user 304 (new employee D) performs their duties, while in Figure 4, the designated space 200 is the workspace where user 301 (department head A) performs their duties. Such designated spaces 200 are simple private rooms for each user 300 (worker) within the office, separated, for example by partitions, and are equipped with desks, chairs, and information terminals.
[0031] In recent years, the introduction of remote work has brought about changes in how workers work, and along with these changes in work styles, the places where each worker performs their duties are also changing. The designated space 200 may be the office, the employee's home (working from home), or a workspace other than the office. If the designated space 200 is the employee's home, the employee takes their information terminal, such as a laptop computer, home and works from home.
[0032] Since the various sensors 6 are conventionally known sensors, a detailed explanation of their functions will be omitted, but they are as follows. The communication standard for communication between each sensor 6 and the information provision system 100 is not particularly limited, and it may be wireless or wired communication. In addition, at least one of the following may be interposed between each sensor 6 and the information provision system 100: an information terminal, gateway, communication adapter, and other servers. Hereinafter, the sensors 6 installed / used in the designated space 200 where the user 300 (target person T1) performs their work may be simply referred to as "the user 300's (or target person T1's) sensors 6". The sensor ID information of each user 300's various sensors 6 may be associated with the user 300's personal information (name, employee number, and email address, etc.) and managed in a database (user information storage unit S14).
[0033] The microphone 61 collects the voice of the subject T1. The microphone 61 can be attached to a partition, wall, or ceiling that constitutes the predetermined space 200. In the illustrated example in Figure 3, the microphone 61 (indicated as sensor 6 in the drawing) is attached to the wall of the predetermined space 200. Alternatively, the microphone 61 may be a headset type microphone that is worn on the head of the subject T1. The microphone 61 may also be a microphone attached to an information terminal. The number of microphones 61 is not particularly limited, and one or more microphones 61 may be installed / used in one predetermined space 200. Since the microphone 61 is easier to install and introduce than other sensors 6, it can be used not only when the predetermined space 200 is a private room in an office, but also when the predetermined space 200 is a home.
[0034] The sound information collected by the microphone 61 is output as an electrical signal (sound signal). The sound signal may be transmitted directly to the information provision system 100. Alternatively, if the microphone 61 is connected to an information terminal or communication adapter, the sound signal may be transmitted to the information provision system 100 via the information terminal or communication adapter. Analog-to-digital conversion processing (AD conversion processing) of sound may be performed on the information provision system 100 side (for example, the information acquisition unit 1), or on the information terminal or communication adapter.
[0035] The motion sensor 62 senses the movement of a subject T1 and outputs information about the presence or absence of the subject T1 in the predetermined space 200 as an electrical signal (human detection signal). The motion sensor 62 is primarily intended for use when the predetermined space 200 is a private room in an office, but it may also be used when the predetermined space 200 is a home if installation and implementation are easy. The motion sensor 62 can be attached to partitions, walls, or ceilings that make up the predetermined space 200 (private room). The number of motion sensors 62 is not particularly limited, and one or more motion sensors 62 can be installed / used in one predetermined space 200.
[0036] The type of motion sensor 62 is not particularly limited; it may be infrared or ultrasonic. The motion sensor signal may be transmitted directly to the information provision system 100. Alternatively, if the motion sensor 62 is connected to an information terminal or communication adapter, the motion sensor signal may be transmitted to the information provision system 100 via the information terminal or communication adapter.
[0037] The position sensor 63 senses the position of the subject T1 and senses the position of the subject T1 within the predetermined space 200, or (if it is an office) the position of the subject T1 within the office outside the predetermined space 200. The position sensor 63 is mainly intended to be used when the predetermined space 200 is a private room in an office, but it may also be used when the predetermined space 200 is a home if it is easy to install and implement.
[0038] More specifically, the location information sensed by the location sensor 63 is information indicating the coordinates of a worker (user 300) in the office. The positioning system, which includes multiple location sensors 63, measures the location information of each of the multiple workers at predetermined time intervals and stores (manages) the time-series data of the location information. The positioning system includes, for example, a positioning server, multiple beacon transmitters (location sensors 63) distributed in a workplace such as a predetermined space 200 (e.g., on the ceiling), and beacon receiving terminals (location sensors 63) carried by each of the multiple workers.
[0039] The beacon receiving terminal measures the signal strength of the beacon signals received from each of the multiple beacon transmitters and transmits strength information indicating the measured signal strength to the positioning server. Based on the received strength information, the positioning server calculates the distance between the beacon receiving terminal and the corresponding beacon transmitter. Based on the distance between the beacon receiving terminal and each of the multiple beacon transmitters, and the position information of each of the multiple beacon transmitters, the positioning server can measure the position information of the beacon receiving terminal (i.e., the worker possessing the beacon receiving terminal) by tripoint positioning. Note that each worker may own a beacon transmitter, and multiple beacon receiving terminals may be distributed in a workplace such as a predetermined space 200 (e.g., on the ceiling). The number of beacon transmitters and beacon receiving terminals is not particularly limited.
[0040] The positioning server transmits the measured location information of the subject T1 as an electrical signal (position signal) to the information provision system 100. The positioning server may be the same server on which the functions of the information provision system 100 are implemented.
[0041] The speech sensor 64 senses the speech of the subject T1 and outputs the speech information of the subject T1 within the predetermined space 200 as an electrical signal (speech signal). It is assumed that the speech sensor 64 has more advanced processing capabilities than the microphone 61 (such as extracting only the frequencies corresponding to human speech from the collected sound and converting the speech into text).
[0042] Since the speech sensor 64 can be easily installed and introduced in the same way as the microphone 61, it can be used not only when the predetermined space 200 is a private room space within an office but also when the predetermined space 200 is a home. The speech sensor 64 can be attached to partitions, walls, ceilings, etc. that make up the predetermined space 200. The speech sensor 64 may also be a bone conduction sensor that is worn on the head or ear of the subject T1 in a headset type or earphone type for use. The speech sensor 64 may be attached to an information terminal.
[0043] The speech signal may be directly transmitted to the information providing system 100. Alternatively, when the speech sensor 64 is communicably connected to an information terminal, a communication adapter, etc., the speech signal may be transmitted to the information providing system 100 via the information terminal, the communication adapter, etc. The number of speech sensors 64 is not particularly limited, and one or more speech sensors 64 can be installed / used for one predetermined space 200.
[0044] In addition, for example, a camera (imaging device) may be installed / used as the sensor 6 within the predetermined space 200. The camera captures the state of the subject T1 who is performing an event activity (business activity) within the predetermined space 200 and transmits the image information as an electrical signal (image signal) to the information providing system 100. However, the use of the camera may not be desirable considering the privacy of the user 300, as all aspects of the user 300's activities will be digitized (visualized) even during business hours. In particular, when the predetermined space 200 is a home, considering the privacy of the user 300, information collection by sensors 6 other than the camera is desirable.
[0045] In the information providing system 100, by using the sensing results of such various sensors 6, it is possible to identify to some extent the action patterns of the subject T1 inside and outside the predetermined space 200 during the period of daily business activities. That is, as shown in FIG. 3, it is possible to identify the periods, order, and number of times of each phase such as the work (concentration) phase, toilet phase, (remote) meeting phase, break phase, and work (relaxed) phase of the subject T1.
[0046] One or more acoustic devices 5, like the sensor 6, are installed / used within a predetermined space 200 where the user 300 conducts event activities (here, business activities). That is, the acoustic device 5 may be installed within the predetermined space 200, or it may be worn by the user 300 on their body and used within the predetermined space 200. The acoustic device 5 includes, for example, a speaker.
[0047] The communication standard for the communication conducted between each acoustic device 5 and the information providing system 100 is not particularly limited, and it may be wireless communication or wired communication. Also, at least one of an information terminal, a gateway, a communication adapter, and other servers, etc. may be interposed between the communication (path) of each acoustic device 5 and the information providing system 100.
[0048] The acoustic device 5 can be attached to a partition, wall, or ceiling, etc. that constitutes the predetermined space 200. In the illustrated example of FIG. 4, the speaker of the acoustic device 5 is attached to the ceiling of the predetermined space 200. Alternatively, the acoustic device 5 may be used while being worn on the head or ears of the user 300 in the form of a headset type or earphone type. The acoustic device 5 may also have the function of the sensor 6 (such as the microphone 61, etc.). The acoustic device 5 may include a speaker attached to an information terminal. The number of acoustic devices 5 is not particularly limited, and one or more microphones 61 can be installed / used for one predetermined space 200. Since the acoustic device 5 is relatively easy to install and introduce, it can be used not only when the predetermined space 200 is an individual room space within an office but also when the predetermined space 200 is a home.
[0049] Hereinafter, the acoustic device 5 installed / used within the predetermined space 200 where the user 300 (target person T1) conducts business may sometimes be simply called, for short, "the acoustic device 5 of the user 300 (or target person T1)". The device ID information of the acoustic device of each user 300 is associated with the personal information (name, employee number, email address, etc.) of that user 300 and is managed on a database (user information storage unit S14).
[0050] The sound device 5 outputs sound based on sound source data corresponding to the situation of the subject T1, in accordance with the sound signal output from the information provision system 100 (sound output unit 4). As an example of use, sound based on sound source data corresponding to the situation of subject T1, user 304 (new employee D), can be output from the sound device 5 of user 301 (department head A), who is user 304's superior (see Figure 4). The above sound can also be output from the sound devices 5 of users 302 to 303, as well as from user 304's own sound device 5.
[0051] The sound signal from the information provision system 100 may be transmitted directly to the sound device 5. Alternatively, if the sound device 5 is connected to an information terminal or communication adapter, the sound signal may be transmitted to the sound device 5 via the information terminal or communication adapter. The signal processing of the sound signal (DA conversion processing and signal amplification processing, etc.) is assumed to be performed on the information provision system 100 side (for example, the sound output unit 4), but it may also be performed in the sound device 5, information terminal, or communication adapter.
[0052] A communication tool is one example of a means by which the information provision system 100 obtains "subject information" used to estimate / evaluate the situation of subject T1. As described above, the communication tool may be an email tool, a chat tool, or an online meeting application tool such as Teams® or Zoom®.
[0053] Dedicated application software for using the communication tool is installed on each user's information terminal (300). The cloud server 7, which manages the communication tool, transmits (provides) information on the communication tool usage status of target person T1 (target person information) to the information provision system 100 in response to a request signal from the information provision system 100. Examples of usage status information include information on sending emails, sending chat messages, participation status and statements in remote meetings, and information on actions and statements in virtual space (between users 300). Note that statements made in remote meetings, etc., may be converted into text and transmitted to the information provision system 100. Communication between the cloud server 7 and the information provision system 100 may be conducted via a wide-area communication network such as the Internet. In addition, at least one of the following may be interposed between the cloud server 7 and the information provision system 100: an information terminal, a gateway, a communication adapter, and other servers.
[0054] (2) Configuration of the Information Provision System Next, a detailed explanation of each function in the information provision system 100 will be given. As described above, the information provision system 100 comprises an information acquisition unit 1, an estimation unit 2, a sound source selection unit 3, a sound output unit 4, an evaluation unit H1, and a storage unit S1 (see Figure 1).
[0055] (2.1) Configuration of the storage unit The storage unit S1 is a storage device that stores (stores) information necessary for various processes implemented in the information provision system 100, as well as computer programs executed by the information provision system 100. The storage unit S1 is, for example, a non-volatile memory.
[0056] The storage unit S1 also includes an organization information storage unit S11, an estimated result storage unit S12, a sound source data storage unit S13, and a user information storage unit S14. Note that the organization information storage unit S11, the estimated result storage unit S12, the sound source data storage unit S13, and the user information storage unit S14 may be configured using separate storage devices.
[0057] The organizational information storage unit S11 includes a database that stores organizational information about one or more organizations to which multiple users 300 belong. Specifically, the organizational information storage unit S11 stores organizational information indicating that users 301 to 304 belong to the same organization (in this case, the Human Resources Department). The organizational information also includes information about the relationships between users 301 to 304 within that organization (job titles such as department head and section chief, and relationships such as superior, subordinate, senior, junior, colleague, and newcomer). The organizational information can be updated as needed in accordance with personnel changes, etc.
[0058] The estimation result storage unit S12 includes a database that stores (stores) the estimation results (from the estimation unit 2, described later) of the subject T1 as stored information. In other words, the estimation result storage unit S12 stores (stores) the estimation results (from the estimation unit 2) of multiple users 300 as stored information. The stored information can be updated each time the estimation process is executed in the estimation unit 2.
[0059] The sound source data storage unit S13 includes a database that pre-stores sound source data for multiple sounds (e.g., natural sounds). Natural sounds include, as described above, bird songs (e.g., birdsong), bird wing sounds, animal footsteps, animal calls (other than birds), insect sounds, stream sounds, wave sounds, rain sounds, waterfall sounds, wind sounds, rustling leaves, campfire sounds, gusts of wind, and thunder. The sound source data can be updated as appropriate in accordance with processing such as version upgrades.
[0060] The user information storage unit S14 includes a database that stores the personal information of the user 300 (name, employee number, email address, etc.). The database of the user information storage unit S14 stores the terminal ID information of each information terminal in association with the personal information of the user 300. The database of the user information storage unit S14 also stores the sensor ID information of each user 300's various sensors 6 in association with the personal information of the user 300. Furthermore, the database of the user information storage unit S14 stores the device ID information of each user 300's sound device 5 in association with the personal information of the user 300.
[0061] Furthermore, the database in the user information storage unit S14 stores the type of sound source data (natural sounds) that each user 300 has selected in advance, and associates it with the personal information of that user 300.
[0062] (2.2) Configuration of the Information Acquisition Unit The Information Acquisition Unit 1 performs information acquisition processing during the event activities of the subject T1 and acquires subject information that includes activity information relating to at least one of the voice spoken by the subject T1, the actions taken by the subject T1, and the characters input by the subject T1. Here, as an example, it is assumed that the subject information includes all of the following: activity information relating to voice (hereinafter simply referred to as voice information), activity information relating to actions (hereinafter simply referred to as action information), and activity information relating to characters (hereinafter simply referred to as character information).
[0063] The information acquisition unit 1 includes a first information acquisition unit 11 and a second information acquisition unit 12.
[0064] The first information acquisition unit 11 has a communication function for communicating with various sensors 6, etc. The first information acquisition unit 11 acquires subject information (hereinafter also referred to as first subject information) including activity information from the various sensors 6. The subject information (first subject information) includes real-space information sensed by one or more sensors 6 as activity information. In other words, since the various sensors 6 are sensors installed / used in the actual space (predetermined space 200) where the subject T1 performs their work, the voice information, behavior information, and text information included in the first subject information are real-space information.
[0065] In the first subject information, the activity information obtained from the microphone 61 is classified as voice information. In other words, the subject information (first subject information) includes voice-related activity information.
[0066] In the first subject information, the activity information obtained from the human presence sensor 62 and the position sensor 63 is classified as behavioral information.
[0067] In the first subject information, the activity information obtained from the speech sensor 64 is based on the voice from subject T1, but since it is converted into text by the speech sensor 64 and transmitted, it is classified as text information.
[0068] The first information acquisition unit 11 can acquire information about the first subject T1 from various sensors 6 at any time while the subject T1 is participating in an event (work activities). In other words, the first subject information can be information about the subject T1 acquired in near real time.
[0069] The first information acquisition unit 11 has an AD conversion function, and if the signal received from the sensor 6 is an analog signal, it converts it to a digital signal as needed.
[0070] The first information acquisition unit 11 inputs the acquired first target information into the first estimation unit 21 (estimation unit 2).
[0071] The second information acquisition unit 12 has a communication function for communicating with the cloud server 7 of the communication tool. The second information acquisition unit 12 calls an API (Application Programming Interface) to establish cooperation (communication) with the cloud server 7 and sends a request signal to the cloud server 7. The request signal can specify the information to be requested from the cloud server 7. In response to the request signal, the second information acquisition unit 12 acquires subject information (hereinafter also referred to as second subject information), including activity information, that is sent from the cloud server 7. The subject information (second subject information) includes virtual space information obtained from a communication tool used by multiple users 300 as activity information. In other words, the second subject information from the cloud server 7 is different from the first subject information from various sensors 6; it is information from a virtual space deployed via the cloud server 7. The voice information, behavior information, and text information included in the second subject information are virtual space information.
[0072] In the second section of the target information, email and chat messages are classified as text information.
[0073] In the second section of participant information, information regarding participation in remote meetings (such as the amount of time spent and the number of remote meetings attended per day) and activities in the virtual space is classified as activity information.
[0074] In the second section of participant information, spoken information within remote meetings and spoken information in virtual spaces are classified as audio information. If spoken information is converted to text on the cloud server 7, it may be classified as text information.
[0075] The second information acquisition unit 12 can periodically acquire second target information from the cloud server 7 at a longer interval than the acquisition cycle for first target information (for example, once a day).
[0076] The second information acquisition unit 12 inputs the acquired second target information into the second estimation unit 22 (estimation unit 2).
[0077] (2.3) Configuration of the estimation unit The estimation unit 2 performs estimation processing to estimate at least one of the emotions and activity levels of subject T1 (in this example, both) based on the subject information.
[0078] Here, "emotions" refer to, for example, subject T1's "happiness," "indifference," "anger," "sadness," "surprise," "displeasure," "fear," and "contempt." "Emotions" may also be defined, for example, based on Russell's Circle of Emotions model. In the following, as an example, we will assume that each of these emotions will be treated with a symbol (hereinafter also called an emotion symbol).
[0079] Furthermore, the "activity level" referred to here is the degree of "motivation" of subject T1, and could be, for example, the degree of willingness to participate in remote meetings or the degree of willingness to engage in work activities. In the following, as an example, we will assume that these activity levels will be handled using symbols (hereinafter also referred to as activity symbols).
[0080] The estimation unit 2 estimates the emotion and activity level of subject T1 using, for example, a trained model (classifier) for which machine learning has been completed. Here, as an example, speaker inference model M1 and emotion inference models M2 and M3 are used as trained models (see Figure 1). The trained models referred to here are assumed to include, for example, models using neural networks or models generated by deep learning using multilayer neural networks. Neural networks may include, for example, CNN (Convolutional Neural Network) or BNN (Bayesian Neural Network). The trained models are realized by implementing the trained neural network on an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or FPGA (Field-Programmable Gate Array). The trained models are not limited to models generated by deep learning. The trained models may also be models generated by support vector machines or decision trees.
[0081] In this embodiment, as an example, the estimation unit 2 is configured to separately perform estimation processing for the first target information and estimation processing for the second target information. Specifically, the estimation unit 2 includes a first estimation unit 21 and a second estimation unit 22.
[0082] The first estimation unit 21 receives the first target person information acquired by the first information acquisition unit 11 and performs estimation processing (first estimation processing) on the first target person information.
[0083] In the first estimation process, speaker inference model M1 and emotion inference model M2 are used. Speaker inference model M1 is used when it is difficult to directly identify the subject T1 using only the first subject information from sensor 6, such as voice information from microphone 61.
[0084] The speaker inference model M1 is a machine learning model that uses the voice data (voiceprint features) of each of the 300 users and their personal identifiers (e.g., employee numbers) as training data (sets). For example, the voice information (activity information) obtained from the microphone 61 may lack information necessary to directly identify the target person T1 based solely on that voice information. Therefore, by inputting the voice information obtained from the microphone 61 as input data into the speaker inference model M1, the speaker inference model M1 outputs the personal identifier of the target person T1 as a classification result based on the voiceprint features in the voice information.
[0085] The emotion inference model M2 is a machine learning model that uses various audio information, behavioral information, and textual information obtainable from various sensors 6, along with emotion symbols and activation symbols, as training data (sets). Audio information, behavioral information, and textual information from various sensors 6 are input to the emotion inference model M2 as input data. Based on the features extracted from the input audio information, behavioral information, and textual information, the emotion inference model M2 outputs the emotion symbols and activation symbols of the subject T1 as classification results. The features may include speech pitch and intonation, the use of specific words in speech or text converted and obtained from communication tools, utterance frequency, and behavioral patterns based on departure status from a predetermined space 200 or changes in location information within the office.
[0086] The first estimation unit 21 may, if necessary, refer to the user information storage unit S14 of the storage unit S1.
[0087] The first estimation unit 21 outputs the personal symbol, emotion symbol, and activity symbol of subject T1, which are estimation results using the speaker inference model M1 and the emotion inference model M2, to the evaluation unit H1. The first estimation unit 21 also stores (accumulates) the personal symbol, emotion symbol, and activity symbol of subject T1 in the estimation result storage unit S12 of the storage unit S1.
[0088] In addition to the information of these symbols, the first estimation unit 21 stores (accumulates) in the estimation result storage unit S12 information such as the time when the first subject information was acquired, and the location where subject T1 was performing their duties (office, home, etc.).
[0089] The second estimation unit 22 receives the second target information acquired by the second information acquisition unit 12 and performs estimation processing (second estimation processing) on the second target information.
[0090] In the second estimation process, only the emotion inference model M3 is used. The speaker inference model is not used in the second estimation process. When using the communication tool, user 300 will enter login information on the information terminal. Therefore, unlike the first target information, the second target information from the cloud server 7 is highly likely to contain information that can directly identify target T1 (personal information such as name and employee number). For this reason, the use of the speaker inference model is unnecessary in the second estimation process, however, as in the first estimation process, the speaker inference model may be used as needed.
[0091] The emotion inference model M3 is a machine learning model that uses various audio information, behavioral information, and textual information obtainable from the cloud server 7, along with emotion symbols and activation symbols, as training data (sets). The audio information, behavioral information, and textual information from the cloud server 7 are input to the emotion inference model M3 as input data. Based on the features extracted from the input audio information, behavioral information, and textual information, the emotion inference model M3 outputs the emotion symbols and activation symbols of the subject T1 as classification results. The features may include speech pitch and intonation, the use of specific words in speech or text converted and obtained from communication tools, utterance frequency, transmission frequency, participation status in remote meetings, and behavioral patterns in virtual space.
[0092] The second estimation unit 22 may, if necessary, refer to the user information storage unit S14 of the storage unit S1.
[0093] The second estimation unit 22 outputs to the evaluation unit H1 the personal code of subject T1 based on personal information extracted from the second subject information, and the emotion code and activity code of subject T1, which are estimation results using the emotion inference model M3. The second estimation unit 22 also stores (accumulates) the personal code, emotion code and activity code of subject T1 in the estimation result storage unit S12 of the storage unit S1.
[0094] In addition to the information of these symbols, the second estimation unit 22 stores (accumulates) in the estimation result storage unit S12 time information regarding the time when the second subject information was acquired, and location information regarding the place where subject T1 was performing their duties (office, home, etc.).
[0095] (2.4) Configuration of the evaluation unit The evaluation unit H1 performs an evaluation process and evaluates the estimation results (personal symbol, emotion symbol, and activity symbol of subject T1) from the estimation unit 2 to generate evaluation information for subject T1. The evaluation unit H1 outputs the generated evaluation information to the sound source selection unit 3.
[0096] Specifically, the evaluation department H1 comprises the individual evaluation department H11, the medium- to long-term evaluation department H12, and the organizational evaluation department H13.
[0097] The individual evaluation unit H11 performs an individual evaluation of subject T1 based on the estimated result of subject T1 at a certain point in time and the accumulated information, by determining the change in the estimated result of subject T1 at a point in time prior to that point, and outputs the result of that individual evaluation as evaluation information.
[0098] Basically, the individual evaluation unit H11 is expected to perform evaluation processing in real time each time it receives an estimation result from the estimation unit 2. However, the frequency of evaluation processing by the individual evaluation unit H11 may be changed as appropriate in response to input from users 300, etc., via an information terminal. For example, the execution frequency may be limited to only a few times a day.
[0099] When the personal evaluation unit H11 receives the latest personal symbol, emotion symbol, and activity symbol of subject T1, which are the estimation results, from the first estimation unit 21 or the second estimation unit 22, it extracts the estimation result of subject T1 at a point in time prior to the time (point in time) of the time information (for example, the most recent point in time) from the stored information of the estimation result storage unit S12. The personal evaluation unit H11 compares the latest emotion symbol and activity symbol with the most recent emotion symbol and activity symbol, and if there is a change of a certain amount or more in any element, it identifies that element and outputs it as evaluation information to the sound source selection unit 3.
[0100] To give a specific example, the individual evaluation unit H11 can make evaluations such as whether the element of happiness (emotions) has increased by a certain amount (changed in a positive direction) or decreased by a certain amount (changed in a negative direction). In addition, the individual evaluation unit H11 can make evaluations such as whether the element of willingness to participate in remote meetings (activity level) has increased by a certain amount (changed in a positive direction) or decreased by a certain amount (changed in a negative direction).
[0101] Here, as an example, the Individual Evaluation Department H11 can also be called the Short-Term Evaluation Department because it seeks to determine the change between the most recent point in time and the most recent point in time.
[0102] The medium- to long-term evaluation unit H12 extracts multiple estimated results for subject T1 within a predetermined period from the accumulated information, evaluates the medium- to long-term trends of subject T1 from the extracted results, and outputs the results of the medium- to long-term evaluation as evaluation information.
[0103] Unlike the individual evaluation unit H11, the medium- to long-term evaluation unit H12 does not perform evaluation processing in real time each time it receives an estimation result from the estimation unit 2. Rather, it is assumed to be performed periodically at relatively long intervals, such as once a week or once a month. Alternatively, the evaluation processing by the medium- to long-term evaluation unit H12 may be performed at any time in response to an execution request from a user 300 or the like via an information terminal.
[0104] The medium- to long-term evaluation unit H12 extracts multiple estimation results for subject T1 within a predetermined period (e.g., the most recent month) from the stored information in the estimation result storage unit S12. The medium- to long-term evaluation unit H12 analyzes (evaluates) the trend of the multiple estimation results within the predetermined period and outputs it as evaluation information to the sound source selection unit 3. The estimation results extracted from the stored information may be estimation results from the first estimation unit 21 or estimation results from the second estimation unit 22.
[0105] To give a specific example, the medium- to long-term evaluation unit H12 may make an assessment such as whether the fluctuations in the element of happiness (emotions) over the past month have been drastic (or mild). Also, the medium- to long-term evaluation unit H12 may make an assessment such as whether the element of willingness to participate in remote meetings (activity level) over the past month has been actively increasing (or chronically decreasing).
[0106] The organizational evaluation unit H13 identifies the organization to which subject T1 belongs based on the estimated results of subject T1, organizational information, and accumulated information, performs a relative evaluation of subject T1 within the organization, and outputs the results of that relative evaluation as evaluation information.
[0107] The execution timing of the organizational evaluation unit H13 is not particularly limited. The organizational evaluation unit H13 may perform evaluation processing in real time whenever it receives estimation results from the estimation unit 2, similar to the individual evaluation unit H11. Furthermore, the execution frequency of evaluation processing by the organizational evaluation unit H13 may be appropriately set and changed in response to input from users 300, etc., via information terminals. For example, the execution frequency may be limited to only a few times a day.
[0108] When the organizational evaluation unit H13 receives the latest personal symbol, emotion symbol, and activity symbol of subject T1, which are estimation results, from the first estimation unit 21 or the second estimation unit 22, for example, it refers to the organizational information in the organizational information storage unit S11 and identifies the organization to which subject T1 belongs. The organizational evaluation unit H13 identifies other users 300 (if subject T1 is user 304, then users 301 to 303 other than user 304) within the identified organization. The organizational evaluation unit H13 extracts the estimation results of other users 300 (users 301 to 303) at a point in time prior to the time (point in time) of the time information (for example, the most recent point in time) from the stored information in the estimation result storage unit S12. The estimation results extracted from the stored information may be the estimation results from the first estimation unit 21 or the estimation results from the second estimation unit 22.
[0109] The organizational evaluation unit H13 then takes into account the emotion symbols and activity symbols of the other 300 users, performs a relative evaluation (organizational evaluation) of the latest emotion symbols and activity symbols of subject T1, and outputs this as evaluation information to the sound source selection unit 3.
[0110] To give a specific example, the Organizational Evaluation Department H13 can evaluate the element of happiness (emotion) by determining whether the element (value) of subject T1 is high or low (superior or inferior) compared to the organization's representative value (mean or median, etc.). Similarly, the Organizational Evaluation Department H13 can evaluate the element of willingness to participate in remote meetings (activity level) by determining whether the element (value) of subject T1 is high or low (superior or inferior) compared to the organization's representative value (mean or median, etc.).
[0111] Furthermore, the Organizational Evaluation Department H13 may, based on an overall assessment, evaluate that subject T1 lacks energy within the organization, or that subject T1 is prone to causing conflict within the organization.
[0112] The organizational evaluation unit H13 may be executed periodically, for example, once a week or once a month, at relatively long intervals, similar to the medium- to long-term evaluation unit H12. Alternatively, the evaluation process by the organizational evaluation unit H13 may be executed at any time in response to execution requests from users 300, etc., via information terminals. When executed periodically or at any time, the organizational evaluation unit H13 may extract not only the estimation results of other users 300 but also the estimation results of the target person T1 from the stored information of the estimation result storage unit S12 and execute the evaluation process. The estimation results extracted from the stored information may be the estimation results from the first estimation unit 21 or the estimation results from the second estimation unit 22.
[0113] (2.5) Configuration of the sound source selection unit The sound source selection unit 3 executes a sound source selection process to select sound source data corresponding to the evaluation information received from the evaluation unit H1 (i.e., evaluation information based on the estimation results by the estimation unit 2) from among multiple sound source data. The sound source selection unit 3 selects the corresponding sound source data by referring to the sound source data storage unit S13 and the user information storage unit S14.
[0114] [Single-shot and continuous sound sources] In this embodiment, as an example, it is assumed that the multiple sound source data includes one or more first sound source data which are highly single-shot sounds, and one or more second sound source data which are highly continuous sounds. However, this is not limited to this, and the multiple sound source data may include only one or more first sound source data, or only one or more second sound source data.
[0115] Highly isolated sounds are sounds related to living things, such as bird calls, bird wing sounds, animal calls (other than birds), animal footsteps, and insect sounds. For each of these categories, such as "bird calls," "bird wing sounds," and "animal calls (other than birds)," a large number of different types of sounds (sound source data) are registered in the database. For example, regarding "bird calls," the database contains the calls of many different types of birds, such as bulbuls, nightingales, chicks, and pigeons.
[0116] On the other hand, highly continuous sounds include environmental sounds such as the sound of a stream, waves, rain, waterfalls, wind, rustling leaves, campfires, gusts of wind, and thunder. For each of these, such as "sound of a stream," "sound of waves," and "sound of rain," numerous types of sounds (sound source data) are registered in the database. For example, regarding "sound of rain," many types of rain sounds are registered in the database, such as the sound of light rain, the sound of heavy rain, and the sound of a storm.
[0117] As shown in Figure 5, the sound source data storage unit S13 includes, for example, a single-song sound source storage area R1 for storing multiple first sound source data and a continuous sound source storage area R2 for storing multiple second sound source data.
[0118] In this embodiment, the sound source selection unit 3 determines (selects) whether to emit a sound with a high degree of individuality or a sound with a high degree of continuity, depending on the type of evaluation information from the evaluation unit H1 (individual evaluation, relative evaluation, short-term evaluation, medium- to long-term evaluation).
[0119] Specifically, the evaluation information from the evaluation unit H1 includes either a first result, which is the result of the individual evaluation of subject T1, or a second result, which is the result of the relative evaluation of subject T1 compared to other users 300. The sound source selection unit 3 selects one from one or more first sound source data if the evaluation information includes the first result, and selects one from one or more second sound source data if the evaluation information includes the second result. Here, "relative evaluation" is assumed to be the evaluation by the organizational evaluation unit H13.
[0120] Furthermore, the evaluation information from the evaluation unit H1 includes a third result, which is the result of a short-term evaluation of subject T1, or a fourth result, which is the result of a medium- to long-term evaluation of subject T1. If the evaluation information includes a third result, the sound source selection unit 3 selects one from one or more first sound source data, and if the evaluation information includes a fourth result, it selects one from one or more second sound source data. Here, "short-term evaluation" is assumed to be an individual evaluation by the individual evaluation unit H11.
[0121] In short, for example, if the evaluation information includes a first result or a third result, the sound source selection unit 3 refers to the single-event sound source storage area R1 and selects one first sound source data from a plurality of first sound source data. Also, for example, if the evaluation information includes a second result or a fourth result, the sound source selection unit 3 refers to the continuous sound source storage area R2 and selects one second sound source data from a plurality of second sound source data.
[0122] By the way, in this embodiment, it is assumed that each user 300 has pre-set the type of sound source data corresponding to themselves in the information provision system 100.
[0123] Each user 300 has pre-selected one preferred type of first sound source data from among several first sound source data. Furthermore, each user 300 has pre-selected one preferred type of second sound source data from among several second sound source data.
[0124] Each user 300 performs an operation input, for example, on an information terminal to select their preferred type of sound source data. In response to this operation input, the information provision system 100 sets (registers, modifies, etc.) the sound source data associated with the user 300 in the database of the user information storage unit S14.
[0125] Figure 2 illustrates, as an example, the first sound source data (birdsong sound source data D1) that each of users 301 to 304 has pre-set.
[0126] In the example in Figure 2, all four users (300) in the Human Resources Department have set the "birdsong" sound source data D1 as the first sound source data for highly isolated sounds. However, the "birdsong" set by the four users (300) is the call of a different type of bird.
[0127] In the example shown in Figure 2, user 301 has pre-configured sound source data D11 of the call of a red-crowned crane, and user 301's personal information and sound source data D11 are associated and registered in the database of the user information storage unit S14. Similarly, user 302 has pre-configured sound source data D12 of the call of a blue rock thrush, and user 302's personal information and sound source data D12 are associated and registered in the database of the user information storage unit S14. Furthermore, user 303 has pre-configured sound source data D13 of the call of a Japanese bush warbler, and user 303's personal information and sound source data D13 are associated and registered in the database of the user information storage unit S14. Finally, user 304 has pre-configured sound source data D14 of the call of a little owl, and user 304's personal information and sound source data D14 are associated and registered in the database of the user information storage unit S14. Although not shown in the diagram, each of the four users 300 has also pre-configured one second sound source data with high continuity, which is registered in the database of the user information storage unit S14.
[0128] The following explains the sound source selection process with specific examples.
[0129] [In the case of individual evaluation (short-term evaluation)] The sound source selection unit 3 receives evaluation information of subject T1 from the individual evaluation unit H11, for example. The evaluation information from the individual evaluation unit H11 includes the result of the individual evaluation (first result) (or it can be said that it includes the third result, which is the result of the short-term evaluation). If there is a change of a certain element in the evaluation information, the sound source selection unit 3 selects the "first sound source data" corresponding to that change in element. At that time, the sound source selection unit 3 first refers to the user information storage unit S14 and identifies the type of first sound source data that subject T1 has set in advance (for user 304, the sound source data D14 of the call of a little owl). The sound source selection unit 3 extracts the first sound source data of the identified type for subject T1 from the single sound source storage area R1 of the sound source data storage unit S13 and outputs the extraction result to the sound output unit 4.
[0130] In this case, if the evaluation information indicates a positive change in the element of subject T1, the sound source selection unit 3 may instruct the sound output unit 4 to adjust the sound of the sound source data to a brighter tone or increase the volume so that the listener can intuitively understand the change. For example, if the evaluation information indicates a positive change in the element of user 304's happiness (emotion), the sound source selection unit 3 may instruct the sound output unit 4 to adjust the sound of the sound source data D14 of the little owl's call to a brighter tone.
[0131] Furthermore, if the evaluation information indicates a negative change in the subject T1's characteristics, the sound source selection unit 3 may instruct the sound output unit 4 to adjust the sound source data, such as changing the tone to a darker tone or reducing the volume, so that the listener can intuitively understand the change. Alternatively, if the characteristic change indicates a negative change, the sound source selection unit 3 may not select sound source data and may not perform sound output.
[0132] Furthermore, each element is assigned a priority, and if the evaluation information contains a mix of positive and negative element changes, the sound of the audio data may be adjusted based on the change of the element with the highest priority.
[0133] Alternatively, if there is a change of a certain element in the evaluation information, the sound source selection unit 3 may output the sound of the selected sound source data, regardless of whether the change is positive or negative.
[0134] Furthermore, if there are no changes in the elements of the evaluation information, the sound source selection unit 3 may stop processing without selecting sound source data, and sound output from the sound output unit 4 may not be performed. However, even if there are no changes in the elements of the evaluation information, sound source data may be selected and output.
[0135] [In the case of medium- to long-term evaluation] For example, suppose the sound source selection unit 3 receives evaluation information of subject T1 from the medium- to long-term evaluation unit H12. The evaluation information from the medium- to long-term evaluation unit H12 includes the results of the medium- to long-term evaluation (fourth result). If the evaluation information indicates that a certain element has a characteristic change trend (such as rapid fluctuations in change, gentle changes, active increases, chronic decreases, etc.), the sound source selection unit 3 selects the "second sound source data" corresponding to that change trend. In this case, first, the sound source selection unit 3 refers to the user information storage unit S14 to identify the type of second sound source data that subject T1 has set in advance. The sound source selection unit 3 extracts the identified type of second sound source data for subject T1 from the continuous sound source storage area R2 of the sound source data storage unit S13 and outputs the extraction result to the sound output unit 4.
[0136] In this case, the sound source selection unit 3 may instruct the sound output unit 4 to adjust the sound source data to a gentler tone, a more intense tone, or change the volume, so that the listener can intuitively understand the characteristic change trends of subject T1. For example, suppose user 304 has pre-set "rain sound" sound source data as the second sound source data with high continuity. If the evaluation information indicates that the element of user 304's happiness (emotion) shows a gentle change trend, the sound source selection unit 3 may instruct the sound output unit 4 to adjust the sound of the rain sound source data to a tone similar to that of a gentle light rain.
[0137] Furthermore, each element is assigned a priority, and if multiple elements in the evaluation information show characteristic change trends, the tone of the sound source data may be adjusted based on the change trends of the elements with higher priority.
[0138] Alternatively, if there is a characteristic change trend in a certain element of the evaluation information, the sound source selection unit 3 may output the sound of the selected sound source data regardless of the content of that change trend.
[0139] Furthermore, if there are no characteristic changing trends in the evaluation information, the sound source selection unit 3 may stop processing without selecting sound source data, and sound output from the sound output unit 4 may not be performed. However, even if there are no characteristic changing trends in the evaluation information, sound source data may be selected and output.
[0140] [In the case of relative evaluation] For example, suppose the sound source selection unit 3 receives evaluation information of subject T1 from the organizational evaluation unit H13. The evaluation information from the organizational evaluation unit H13 includes the result of a relative evaluation (second result). If the evaluation information indicates superiority or inferiority (high / low) of a certain element relative to the organization's representative value (mean or median, etc.), the sound source selection unit 3 selects the "second sound source data" corresponding to that superiority or inferiority. At that time, first the sound source selection unit 3 refers to the user information storage unit S14 to identify the type of second sound source data that subject T1 has set in advance. The sound source selection unit 3 extracts the second sound source data of the identified type for subject T1 from the continuous sound source storage area R2 of the sound source data storage unit S13 and outputs the extraction result to the sound output unit 4.
[0141] In this case, the sound source selection unit 3 may instruct the sound output unit 4 to adjust the sound source data to a gentler tone, a more intense tone, or change the volume, so that the listener can intuitively understand the superiority or inferiority of subject T1 within the organization. For example, suppose user 304 has pre-set "rain sound" sound source data as the second sound source data with high continuity. If the evaluation information indicates that user 304's happiness (emotions) is "excellent (good)" within the Human Resources Department (organization), the sound source selection unit 3 may instruct the sound output unit 4 to adjust the sound of the rain sound source data to a tone similar to a gentle light rain.
[0142] Furthermore, each element is assigned a priority, and if the evaluation information contains elements of superiority and inferiority, the sound of the audio data may be adjusted based on the element with the higher priority.
[0143] Alternatively, if even one element in the evaluation information is "excellent (good)", the sound source selection unit 3 may output the sound of the selected sound source data.
[0144] [Identifying the output destination] When the sound source selection unit 3 outputs the extraction result (selected sound source data) to the sound output unit 4, it identifies the output destination (sound device 5) of the sound based on the personal code of the subject T1.
[0145] For example, the sound source selection unit 3 refers to the organizational information storage unit S11 and, if the target person T1 is user 304 (new employee D), identifies user 301 (department head A) and / or user 302 (section chief B), who are user 304's superiors within the same organization. Then, the sound source selection unit 3 refers to the user information storage unit S14 and identifies the device ID information of the sound device 5 for user 301 and / or user 302.
[0146] The sound source selection unit 3 outputs output information to the sound output unit 4, which includes the extraction result (selected sound source data) and the device ID information of the output destination.
[0147] It is not mandatory to limit the sound output destination to the supervisor's sound device 5. The sound output destination may include the sound device 5 of user 303, who is user 304's senior, instead of the supervisor's sound device 5, or it may include the sound devices 5 of all members within the same Human Resources Department (organization).
[0148] The sound output destination may also include the sound device 5 of the user 304, who is the subject T1. In that case, the subject T1 can also know how the information provision system 100 estimates / evaluates their own situation.
[0149] The process of identifying the output destination of this sound may also be performed in the sound output unit 4.
[0150] (2.6) Configuration of the sound output unit The sound output unit 4 performs sound output processing and provides information by outputting sound from the sound device 5 based on the extraction result (sound source data selected by the sound source selection unit 3).
[0151] Specifically, the sound output unit 4 receives output information from the sound source selection unit 3, which includes the extraction result (selected sound source data) and the device ID information of the output destination. Based on the output information, the sound output unit 4 performs decoding processing on compressed sound source data if necessary, according to the data format of the sound source data stored in the sound source data storage unit S13, and also performs digital-to-analog conversion processing (DA conversion processing) of the sound. Decoding processing and DA conversion processing may be performed individually for the first sound source data and the second sound source data, respectively.
[0152] Furthermore, the sound output unit 4 performs signal addition processing on the output sound if necessary. The information provision system 100 can also output a sound corresponding to the target person T1 by superimposing a highly continuous sound (sound from the second sound source data) with a highly single-event sound (sound from the first sound source data). For example, while the sound output unit 4 is outputting "rain sounds" based on a medium- to long-term evaluation or relative evaluation as the sound of user 304 from the sound device 5, it can also superimpose and output "little owl calls" based on an individual evaluation (short-term evaluation). In such a case, the sound output unit 4 adds the signal of the "little owl calls" to the signal of the "rain sounds" if necessary.
[0153] Furthermore, the sound output unit 4 performs signal amplification processing and adjusts the sound output level. Also, if the output information includes adjustment instructions, the sound output unit 4 performs adjustment processing such as adjusting the tone of the output sound according to the adjustment instructions.
[0154] The sound output unit 4 then transmits sound (sound signal) to the sound device 5, which is the output destination for the device ID information, causing the sound device 5 to output sound.
[0155] Highly continuous sounds (sound from the second sound source data) are assumed to be repeatedly output from the sound device 5 for a relatively long period (for example, constantly throughout a day's work activities). On the other hand, highly isolated sounds (sound from the first sound source data) are assumed to be output from the sound device 5 for a relatively short period (for example, for several tens of seconds during a day's work activities). However, since highly isolated sounds (sound from the first sound source data) are sounds based on short-term evaluations, they may be output relatively frequently throughout a day's work activities.
[0156] (3) Operation Description (3.1) Operation Example 1 Below, Operation Example 1 of the information provision system 100 will be explained with reference to the sequence diagram shown in Figure 6. Note that the sequence diagram shown in Figure 6 is merely one example of the operation flow related to the information provision system 100, and the order of processing may be changed as appropriate, and processing may be added or omitted as appropriate.
[0157] Operation Example 1 shows an example of operation in conjunction with sensor 6.
[0158] First, various sensors 6 (microphones 61, etc.) in the designated space 200 of the subject T1 perform sensing (step ST1). The sensors 6 transmit the sensing result information as an electrical signal to the information provision system 100 (step ST2).
[0159] When the information provision system 100 receives the sensing result information, it executes a first information acquisition process (information acquisition process) to acquire first target person information including activity information (step ST3). The information provision system 100 also executes a first estimation process (estimation process) based on the first target person information and uses the speaker inference model M1 and the emotion inference model M2 to determine the personal symbol, emotion symbol, and activity symbol of target person T1 (step ST4).
[0160] The information provision system 100 performs an evaluation process based on the subject T1's personal symbol, emotion symbol, and activity symbol to generate evaluation information for subject T1 (step ST5). The evaluation process may include individual evaluation (short-term evaluation), medium- to long-term evaluation, or organizational evaluation (relative evaluation).
[0161] The information provision system 100 performs a sound source selection process based on the evaluation information and selects the corresponding sound source data from among multiple sound source data (step ST6). The information provision system 100 also identifies the device ID information of the sound device 5 that will be the output destination (step ST7).
[0162] The information provision system 100 performs sound output processing (decoding, DA conversion, signal amplification, and adjustment) (step ST8) and transmits sound (sound signal) based on the selected sound source data to the sound device 5 (step ST9). As a result, the sound device 5 outputs sound (step ST10).
[0163] (3.2) Operation Example 2 Below, Operation Example 2 of the information provision system 100 will be explained with reference to the sequence diagram shown in Figure 7. Note that the sequence diagram shown in Figure 7 is merely one example of the operation flow for the information provision system 100, and the order of processing may be changed as appropriate, and processing may be added or omitted as appropriate.
[0164] Operation Example 2 shows an example of operation in conjunction with the cloud server 7 of the communication tool.
[0165] Cloud server 7 manages the usage status of communication tools for multiple users (300).
[0166] The information provision system 100 calls an API to establish cooperation (communication) with the cloud server 7 and also sends a request signal to the cloud server 7 (step ST11).
[0167] In response to a request signal from the information provision system 100, the cloud server 7 transmits information about the communication tool usage status of the subject T1 to the information provision system 100 (step ST12).
[0168] When the information provision system 100 receives information on usage status, it executes a second information acquisition process (information acquisition process) to acquire second target person information, including activity information (step ST13). Based on the second target person information, the information provision system 100 also executes a second estimation process (estimation process) to determine the personal code, emotion code, and activity code of target person T1 using the emotion inference model M3 (step ST14).
[0169] The information provision system 100 performs an evaluation process based on the subject T1's personal symbol, emotion symbol, and activity symbol to generate evaluation information for subject T1 (step ST15). The evaluation process may include individual evaluation (short-term evaluation), medium- to long-term evaluation, or organizational evaluation (relative evaluation).
[0170] The information provision system 100 performs a sound source selection process based on the evaluation information and selects the corresponding sound source data from among multiple sound source data (step ST16). The information provision system 100 also identifies the device ID information of the sound device 5 that will be the output destination (step ST17).
[0171] The information provision system 100 performs sound output processing (decoding, DA conversion, signal amplification, and adjustment) (step ST18) and transmits sound (sound signal) based on the selected sound source data to the sound device 5 (step ST19). As a result, the sound device 5 outputs sound (step ST20).
[0172] (4) Advantages According to the information provision system 100 of this embodiment, the situation of the target person T1 (e.g., user 304) performing work activities can be efficiently transmitted from the sound device 5 of the user 300 (e.g., users 301 and / or 302) as sound (e.g., natural sounds) of the corresponding sound source data. In short, efficient communication of information can be made by "sound," for example, from a subordinate to a superior. In particular, since the information is not provided by direct verbal communication from the target person T1 himself, the possibility of the receiving users 301 and / or 302 having their concentration disrupted and their activities hindered is lower compared to when the communication is done directly by verbal communication. As a result, users 301 and / or 302 can become aware of the situation of the target person T1 through natural awareness, and their actions / ideas are stimulated. Consequently, the information provision system 100 has the advantage of making it easier to provide information that is in harmony with the situation.
[0173] In particular, if the 300 users are workers, providing information that is integrated into the work environment could lead to improvements in productivity, sales, customer satisfaction, and employee retention rates within the organization.
[0174] Furthermore, the reliability of the estimation results and evaluation information by the estimation unit 2 is improved by including real-space information sensed by the sensor 6 in the subject information. In particular, the reliability of the estimation results and evaluation information by the estimation unit 2 is further improved by including activity information related to the voice of subject T1 collected by the microphone 61 in the subject information.
[0175] Furthermore, by including virtual space information obtained from communication tools used by user 300 in the target information, the reliability of the estimation results and evaluation information by the estimation unit 2 is improved.
[0176] (5) Modifications The following are examples of modifications. Each modification described below can be applied in appropriate combination with the above embodiment or other modifications.
[0177] Functions similar to those of the information provision system 100 according to the above embodiment may be embodied in an information provision method, a computer program, or a non-temporary recording medium on which a computer program is recorded.
[0178] The information provision system 100 in this disclosure includes a computer system. The computer system mainly consists of a processor and memory as hardware. The functions of the information provision system 100 in this disclosure are realized by the processor executing a program recorded in the memory of the computer system. The program may be pre-recorded in the memory of the computer system, provided via a telecommunications line, or provided on a non-temporary recording medium such as a memory card, optical disk, or hard disk drive that can be read by the computer system. The processor of the computer system consists of one or more electronic circuits including semiconductor integrated circuits (ICs) or large-scale integrated circuits (LSIs). The integrated circuits such as ICs and LSIs referred to here are named differently depending on the degree of integration, and include integrated circuits called system LSIs, VLSIs (Very Large Scale Integration), or ULSIs (Ultra Large Scale Integration). Furthermore, FPGAs (Field-Programmable Gate Arrays) that are programmed after the manufacture of the LSI, or logic devices that allow for the reconfiguration of junction relationships or circuit compartments within the LSI, can also be used as processors. Multiple electronic circuits may be integrated onto a single chip or distributed across multiple chips. Multiple chips may be integrated onto a single device or distributed across multiple devices. The computer system referred to here includes a microcontroller having one or more processors and one or more memories. Therefore, the microcontroller also consists of one or more electronic circuits, including semiconductor integrated circuits or large-scale integrated circuits.
[0179] Furthermore, it is not essential that the multiple functions of the information provision system 100 be consolidated within a single housing. For example, the components of the information provision system 100 may be distributed across multiple housings.
[0180] Conversely, multiple functions of the information provision system 100 may be consolidated within a single housing. Furthermore, at least some of the functions of the information provision system 100, for example, some of the functions of the information provision system 100, may be implemented by the cloud (cloud computing), etc.
[0181] In the above embodiment, the information provision system 100 is linked to both the sensor 6 and the cloud server 7. However, this is not limited to the above, and the information provision system 100 may be linked to only the sensor 6, or to only the cloud server 7.
[0182] In the above embodiment, real-space information sensed by sensor 6 and virtual-space information obtained from a communication tool are handled separately. In other words, in the above embodiment, the estimation process is divided into a first estimation unit 21 that performs estimation processing on real-space information and a second estimation unit 22 that performs estimation processing on virtual-space information, and estimation processing is performed separately. However, this is not limited to this, and the subject information may include, as activity information, real-space information sensed by one or more sensors 6 installed in a predetermined space 200 where the subject T1 performs event activities, and virtual-space information obtained from a communication tool used by multiple users 300. The estimation unit 2 may integrate the real-space information and the virtual-space information to estimate the subject T1. For example, if there is a difference between the estimation result based on real-space information and the estimation result based on virtual-space information for the same subject T1, the estimation result based on real-space information may be given priority in the evaluation.
[0183] The subject information may include biometric information in addition to activity information. Biometric information may be obtained using skin temperature measurement with a non-contact thermal camera. In this case, the biometric information may include, for example, the skin temperature of user 300. The thermal camera communicates with the information provision system 100 and transmits the measured biometric information (skin temperature). Based on the activity information and biometric information, the information provision system 100 estimates at least one of the subject T1's emotions and activity level.
[0184] (Summary) Based on the embodiments described above, the following embodiments are disclosed.
[0185] The first embodiment of the information provision system (100) provides information about the status of a subject (T1), which is any user (300) among a plurality of users (300). The information provision system (100) comprises an information acquisition unit (1), an estimation unit (2), a sound source selection unit (3), and a sound output unit (4). The information acquisition unit (1) acquires subject information, including activity information relating to at least one of the voices spoken by the subject (T1), the actions taken by the subject (T1), and the characters input by the subject (T1) during the subject (T1)'s event activities. The estimation unit (2) estimates at least one of the subject (T1)'s emotions and level of activity based on the subject information. The sound source selection unit (3) selects sound source data corresponding to the evaluation information based on the estimation result by the estimation unit (2) from among a plurality of sound source data. The sound output unit (4) provides information by outputting sound from the sound device (5) based on the sound source data selected by the sound source selection unit (3).
[0186] According to the above configuration, the subject's (T1) situation is output as sound based on corresponding sound source data, which has the advantage of making it easier to provide information that is more in line with the context compared to providing information through language.
[0187] With respect to the information provision system (100) relating to the second embodiment, in the first embodiment, the subject information includes, as activity information, real-space information sensed by one or more sensors (6) installed / used within a predetermined space (200) where the subject (T1) performs event activities.
[0188] According to the above embodiment, the reliability of the estimation results and evaluation information by the estimation unit (2) is improved by including real-space information in the subject information.
[0189] With respect to the information provision system (100) according to the third embodiment, in the second embodiment, the subject information includes activity information related to voice. One or more sensors (6) include a microphone (61) for collecting voice.
[0190] According to the above embodiment, the reliability of the estimation results and evaluation information by the estimation unit (2) is improved by including activity information related to the voice of the subject (T1) collected by the microphone (61) in the subject information.
[0191] With respect to the information provision system (100) relating to the fourth aspect, in any one of the first to third aspects, the target person information includes virtual space information obtained from a communication tool used by multiple users (300) as activity information.
[0192] According to the above embodiment, the reliability of the estimation results and evaluation information by the estimation unit (2) is improved by including virtual space information in the subject information.
[0193] Regarding the information provision system (100) relating to the fifth aspect, in any one of the first to fourth aspects, the subject information includes, as activity information, real-space information sensed by one or more sensors (6) installed in a predetermined space (200) where the subject (T1) performs event activities, and virtual-space information obtained from communication tools used by multiple users (300). The estimation unit (2) integrates the real-space information and the virtual-space information to estimate the subject (T1).
[0194] According to the above embodiment, the estimation unit (2) integrates real-space information and virtual-space information to estimate the subject (T1), thereby improving the reliability of the estimation results and evaluation information by the estimation unit (2).
[0195] The information provision system (100) according to the sixth embodiment further comprises, in any one of the first to fifth embodiments, an organization information storage unit (S11), an estimation result storage unit (S12), and an organization evaluation unit (H13). The organization information storage unit (S11) stores organization information relating to one or more organizations to which multiple users (300) belong. The estimation result storage unit (S12) stores the estimation results of multiple users (300) as stored information. The organization evaluation unit (H13) identifies the organization to which the subject (T1) belongs based on the estimation result of the subject (T1), the organization information, and the stored information, performs a relative evaluation of the subject (T1) within the organization, and outputs the result of the relative evaluation as evaluation information.
[0196] According to the above configuration, since sound based on sound source data corresponding to the relative evaluation results of the subject (T1) within the organization is output, it is possible to provide information about the relative evaluation results of the subject (T1) in a way that is in harmony with the context.
[0197] The information provision system (100) according to the seventh embodiment further comprises an estimation result storage unit (S12) and a personal evaluation unit (H11) in any one of the first to sixth embodiments. The estimation result storage unit (S12) stores the estimation result of the subject (T1) as stored information. The personal evaluation unit (H11) determines the change in the estimation result of the subject (T1) at a point in time and the stored information, performs a personal evaluation of the subject (T1) based on the estimation result of the subject (T1) at a point in time prior to a certain point in time, and outputs the result of the personal evaluation as evaluation information.
[0198] According to the above configuration, since sound based on audio data corresponding to the results of the subject's (T1) personal evaluation is output, it is possible to provide information about the results of the subject's (T1) personal evaluation in a way that is in harmony with the context.
[0199] The information provision system (100) according to the eighth aspect further comprises an estimation result storage unit (S12) and a medium- to long-term evaluation unit (H12) in any one of the first to seventh aspects. The estimation result storage unit (S12) stores the estimation results of the subject (T1) as stored information. The medium- to long-term evaluation unit (H12) extracts multiple estimation results of the subject (T1) within a predetermined period from the stored information, evaluates the medium- to long-term trends of the subject (T1) from the extracted results, and outputs the results of the medium- to long-term evaluation as evaluation information.
[0200] According to the above configuration, since sound based on sound source data corresponding to the medium- to long-term evaluation results of the subject (T1) is output, it is possible to provide information about the medium- to long-term evaluation results of the subject (T1) in a way that is in harmony with the context.
[0201] Regarding the information provision system (100) relating to the ninth aspect, in any one of the first to eighth aspects, the multiple sound source data includes one or more first sound source data which are highly isolated sounds, and one or more second sound source data which are highly continuous sounds. The evaluation information includes a first result which is the result of the subject's (T1) personal evaluation, or a second result which is the result of the subject's (T1) relative evaluation with other users (300). The sound source selection unit (3) selects one from the one or more first sound source data if the evaluation information includes a first result, and selects one from the one or more second sound source data if the evaluation information includes a second result.
[0202] According to the above configuration, a user (300) listening to the output sound can intuitively determine whether the output sound is a single, continuous sound or a continuous sound, in order to determine whether the information provided is about the individual evaluation of the subject (T1) or the relative evaluation of the subject (T1).
[0203] Regarding the information provision system (100) relating to the tenth aspect, in any one of the first to ninth aspects, the multiple sound source data includes one or more first sound source data which are highly isolated sounds, and one or more second sound source data which are highly continuous sounds. The evaluation information includes a third result which is the result of a short-term evaluation of the subject (T1), or a fourth result which is the result of a medium- to long-term evaluation of the subject (T1). The sound source selection unit (3) selects one from the one or more first sound source data if the evaluation information includes the third result, and selects one from the one or more second sound source data if the evaluation information includes the fourth result.
[0204] According to the above configuration, a user (300) listening to the output sound can intuitively determine whether the output sound is a short-term evaluation result of the subject (T1) or a medium- to long-term evaluation result of the subject (T1), based on whether the output sound is a single, isolated sound or a continuous sound.
[0205] The information provision method relating to the eleventh aspect is an information provision method executed by one or more processors that provides information about the situation of a subject (T1), which is any user (300) among a plurality of users (300). The information provision method includes an information acquisition process, an estimation process, a sound source selection process, and a sound output process. In the information acquisition process, subject information is acquired, which includes activity information relating to at least one of the voices spoken by the subject (T1), the actions taken by the subject (T1), and the characters input by the subject (T1) during the subject (T1)'s event activities. In the estimation process, at least one of the subject (T1)'s emotions and activity level is estimated based on the subject information. In the sound source selection process, sound source data corresponding to the evaluation information based on the estimation results from the estimation process is selected from a plurality of sound source data. In the sound output process, information is provided by outputting sound based on the sound source data selected in the sound source selection process from an audio device.
[0206] According to the above configuration, it is possible to provide an information provision method that makes it easier to provide information that is in harmony with the context.
[0207] The program according to the twelfth embodiment is a program that causes one or more processors to execute the information provision method according to the eleventh embodiment.
[0208] According to the above configuration, it is possible to provide functions that make it easier to provide information that is in harmony with the context.
[0209] 100 Information Provision System 1 Information Acquisition Unit 2 Estimation Unit 3 Sound Source Selection Unit 4 Sound Output Unit 5 Acoustic Device 6 Sensor 61 Microphone 200 Designated Space 300 Users H11 Individual Evaluation Unit H12 Medium- to Long-Term Evaluation Unit H13 Organizational Evaluation Unit S11 Organizational Information Storage Unit S12 Estimation Result Storage Unit T1 Target Person
Claims
An information provision system that provides information about the status of a target individual, who is an arbitrary user among multiple users, An information acquisition unit acquires subject information including activity information relating to at least one of the voices spoken by the subject, the actions taken by the subject, and the characters entered by the subject during the subject's event activities. An estimation unit that estimates at least one of the subject's emotions and activity level based on the subject's information, A sound source selection unit selects sound source data from among multiple sound source data that corresponds to evaluation information based on the estimation results of the estimation unit, The sound output unit provides information by outputting sound from an audio device based on the sound source data selected by the sound source selection unit, Equipped with, Information provision system. The aforementioned subject information includes, as activity information, real-space information sensed by one or more sensors installed / used within a predetermined space where the subject performs the event activity. The information provision system according to claim 1. The aforementioned subject information includes the aforementioned activity information related to the audio, The one or more sensors include a microphone for collecting the sound. The information provision system according to claim 2. The aforementioned subject information includes, as activity information, virtual space information obtained from communication tools used by the multiple users. An information provision system according to any one of claims 1 to 3. The aforementioned subject information includes, as activity information, real-space information sensed by one or more sensors installed in a predetermined space where the subject performs the event activity, and virtual-space information obtained from communication tools used by the multiple users. The estimation unit integrates the real-space information and the virtual-space information to estimate the subject. An information provision system according to any one of claims 1 to 4. Organization information storage unit that stores organizational information relating to one or more organizations to which the aforementioned multiple users belong, An estimation result storage unit that stores the estimation results of the aforementioned multiple users as stored information, An organizational evaluation unit identifies the organization to which the subject belongs based on the estimation results of the subject, the organizational information, and the stored information, performs a relative evaluation of the subject within the organization, and outputs the results of that relative evaluation as evaluation information. It also has, An information provision system according to any one of claims 1 to 5. An estimation result storage unit that stores the estimation results of the subject as stored information, A personal evaluation unit that, based on the estimated result of the subject at a certain point in time and the accumulated information, determines the change in the estimated result of the subject at a point in time prior to that point in time, performs a personal evaluation of the subject, and outputs the result of that personal evaluation as evaluation information. It also has, An information provision system according to any one of claims 1 to 6. An estimation result storage unit that stores the estimation results of the subject as stored information, A medium- to long-term evaluation unit extracts multiple estimation results for the subject within a predetermined period from the accumulated information, evaluates the subject's medium- to long-term trends from the extracted results, and outputs the results of the medium- to long-term evaluation as evaluation information. It also has, The information provision system according to any one of claims 1 to 7. The aforementioned plurality of sound source data includes one or more first sound source data which are highly isolated sounds, and one or more second sound source data which are highly continuous sounds. The aforementioned evaluation information includes a first result, which is the result of the individual evaluation of the subject, or a second result, which is the result of the relative evaluation of the subject compared to other users. The aforementioned sound source selection unit is If the evaluation information includes the first result, select one of the one or more first sound source data, If the evaluation information includes the second result, select one from the one or more second sound source data. An information provision system according to any one of claims 1 to 8. The aforementioned plurality of sound source data includes one or more first sound source data which are highly isolated sounds, and one or more second sound source data which are highly continuous sounds. The aforementioned evaluation information includes a third result, which is the result of a short-term evaluation of the subject, or a fourth result, which is the result of a medium- to long-term evaluation of the subject. The aforementioned sound source selection unit is If the evaluation information includes the third result, select one from the one or more first sound source data. If the evaluation information includes the fourth result, select one of the one or more second sound source data. An information provision system according to any one of claims 1 to 9. An information provision method that is executed by one or more processors and provides information about the status of a target person who is any user among multiple users, Information acquisition process to acquire subject information including activity information relating to at least one of the voice spoken by the subject, the actions taken by the subject, and the characters entered by the subject during the subject's event activities, An estimation process that estimates at least one of the subject's emotions and activity level based on the subject's information, A sound source selection process that selects sound source data corresponding to evaluation information based on the estimation results from the estimation process from among multiple sound source data, The sound output process provides the information by outputting sound from the sound device based on the sound source data selected in the sound source selection process, including, Information provision method. A program that causes one or more processors to execute the information provision method described in claim 11.