Voice output tool, voice control device, voice control method, and application

The audio output device adjusts audio based on user touch and spoken audio, improving user satisfaction and interaction by matching audio output to touch and voice input.

JP2025121700AInactive Publication Date: 2025-08-20SIDEPEAK CO LTD

Patent Information

Application Number
JP2024017331
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-07
Publication Date
2025-08-20
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing audio output devices do not adequately adjust audio output to correspond to the pressure state when a user touches the device and the audio emitted by the user.

Method used

An audio output device equipped with a pressure sensor and a control unit that determines audio output based on the user's contact state and uttered audio, using a system that includes a microphone, speaker, and a control unit to adjust audio accordingly.

Benefits of technology

The audio output is adjusted to match the user's touch state and uttered audio, enhancing user satisfaction and preventing discomfort, allowing for a more engaging interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025121700000001_ABST
    Figure 2025121700000001_ABST
Patent Text Reader

Abstract

To provide a voice output tool which can produce the content of a voice to be output from the voice output tool, in accordance with a pressure state in which a user touches the voice output tool and a voice emitted by the user.SOLUTION: A voice output tool 11 includes: a main body; a speaker 21 which is provided on the main body to output a voice to a user; a pressure sensor 19 which is provided on the main body to detect pressure at which the user has contacted the main body; and a control part 15 which is provided on the main body to determine an output voice to be output from the speaker 21. The voice output tool 11 is configured so that the control part 15 executes: first processing for acquiring voice information emitted by a user; second processing for acquiring a contact state of the user with respect to the main body, which is determined from pressure detected by the pressure sensor 19; and third processing for determining the content of the output voice to be output from the speaker 21 in accordance with the contact state of the user with respect to the main body and the voice information emitted by the user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an audio output tool, an audio control device, an audio control method, and an application that can output audio according to a user's contact state. [Background technology]

[0002] Patent Document 1 describes an information processing device (audio control device). The information processing device includes a control unit that controls sound output from a speaker provided in the device based on the detected state of the device (audio output implement). The control unit continuously changes the output mode of the synthesized sound that the device can output from the speaker in a normal state, depending on the amount of change in the state. Patent Document 1 describes a robot as the device and a pressure sensor that detects the state of the robot. The pressure sensor described in Patent Document 1 detects contact by a user, such as pressing with a finger, for example. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] International Publication No. 2020-129422 Summary of the Invention [Problem to be solved by the invention]

[0004] The inventor of the present application recognized the problem that the voice control device described in Patent Document 1 has room for improvement in that the voice emitted from the voice output device corresponds to the pressure state when the user touches the voice output device, and is not a conversation corresponding to the voice emitted by the user.

[0005] An object of the present invention is to provide an audio output device, an audio control device, an audio control method, and an application that can adjust the audio output from the audio output device to correspond to the pressure state with which the user touches the audio output device and the audio emitted by the user. [Means for solving the problem]

[0006] This embodiment describes an audio output device having a main body, a speaker provided on the main body and outputting audio to a user, a pressure sensor provided on the main body and detecting the pressure with which the user contacts the main body, and a control unit provided on the main body and determining the output audio to be output from the speaker, wherein the control unit performs a first process of acquiring audio information uttered by the user, a second process of acquiring the user's contact state with the main body as determined from the pressure detected by the pressure sensor, and a third process of determining the content of the output audio to be output from the speaker in accordance with the user's contact state with the main body and the audio information uttered by the user. [Effects of the Invention]

[0007] According to this embodiment, the sound output from the sound output tool can be adjusted to correspond to the pressure state with which the user touches the sound output tool and the sound uttered by the user. [Brief explanation of the drawings]

[0008] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a voice control system including a voice output device, a user terminal, and a management server. [Figure 2] FIG. 2 is a schematic diagram illustrating a configuration of a sound output tool. [Figure 3] FIG. 2 is a schematic diagram showing the configuration of a user terminal. [Figure 4] FIG. 2 is a schematic diagram illustrating a configuration of a management server. [Figure 5] 1 is a flowchart showing an example of a voice control method performed in the voice control system. [Figure 6] 10 is a diagram showing an example of pressure information and user voice information used in the voice control system. [Figure 7] 1 is a diagram showing an example of video information and emotion information used in a voice control system. [Figure 8]10 is a diagram showing the relationship between the amount of change in pressure applied to the voice output product and the manner in which the voice output product is touched. DETAILED DESCRIPTION OF THE INVENTION

[0009] (Voice control system explanation) The audio control system is a system that controls audio output from a speaker provided in an audio output tool. The audio control system controls the audio output from the speaker based on the pressure state applied to the audio output tool, the audio of a user touching the audio output tool, etc. Some embodiments included in the audio control system are described with reference to the drawings.

[0010] FIG. 1 shows a voice control system 10 of this embodiment. The voice control system 10 has a voice output device 11, a user terminal 12, and a management server 13. The user terminal 12 and the voice output device 11 can communicate with each other via a network 14. The user terminal 12 and the management server 13 can communicate with each other via the network 14. The network 14 is configured using at least one of wireless communication and wired communication. For convenience, FIG. 1 shows only one voice output device 11 and one user terminal 12. The management server 13 can communicate with multiple user terminals 12. Furthermore, the single user terminal 12 can communicate with multiple voice output devices 11.

[0011] (Explanation of audio output device) The sound output device 11 includes a plurality of different types, such as a cushion, a huggable pillow, a stuffed animal, a toy, a doll, and a robot. The sound output device 11 has a main body 50 shown in FIG. 1. The main body 50 is provided with a control unit 15, a memory unit 16, a communication unit 17, a battery 18, a pressure sensor 19, a microphone 20, a speaker 21, and a camera 22, all of which are shown in FIG. 2. The electric and electronic components provided in the main body 50 function using power supplied from the battery 18. The control unit 15 is connected to the memory unit 16, the communication unit 17, the battery 18, the speaker 21, the microphone 20, the pressure sensor 19, and the camera 22.

[0012] The control unit 15 has an input port, an output port, an arithmetic processing circuit, a main memory circuit, etc. The communication unit 17 is a device that receives and transmits radio waves, electricity, and light to communicate with the outside. The communication unit 17 communicates with the user terminal 12, for example, by short-range waves. The connectable range of short-range waves is, for example, within a radius of 5 to 15 meters from the sound output tool 11. The memory unit 16 stores non-transitory applications, non-transitory programs, information, data, etc. used for processing performed by the control unit 15, such as control and judgment. The information stored in the memory unit 16 includes second identification information assigned to each sound output tool 11. The meaning of the second identification information will be described later.

[0013] If the sound output device 11 is, for example, a cushion, a huggable pillow, or a stuffed toy, the material constituting the main body 50 can be a material that elastically deforms when touched by the user's hand, such as cloth, urethane, sponge, or a composite material thereof. If the sound output device 11 is, for example, a robot, the material of the main body 50 can be synthetic resin, metal, or the like. Note that even when the material of the main body 50 is synthetic resin, metal, or the like, a portion of the main body 50 can be made of a material that elastically deforms when touched by the user's hand, such as cloth, urethane, sponge, synthetic rubber, or a composite material thereof.

[0014] A plurality of pressure sensors (pressure-sensitive devices) 19 are provided inside the main body 50 of the sound output device 11. The plurality of pressure sensors 19 are arranged at different positions along the surface of the main body 50 of the sound output device 11. The pressure sensors 19 may be, for example, capacitive, resistive, or optical. The pressure sensors 19 each have a surface portion 19A and a deep portion 19B. The "surface portion" refers to a portion of the pressure sensor 19 that is shallower from the surface of the main body 50 that the user touches with their hand than the "deep portion." For this reason, for example, when a user strokes the surface of the main body 50 of the sound output device 11 with their hand, the amount of change in pressure received by the surface portion 19A is greater than the amount of change in pressure received by the deep portion 19B. The plurality of pressure sensors 19 each output a signal in accordance with the amount of change in pressure received by the surface portion 19A and the deep portion 19B.

[0015] The microphone 20 collects the voice uttered by the user and outputs a signal corresponding to the voice. The speaker 21 outputs the output voice acquired by the control unit 15. The output voice includes spoken words, music, onomatopoeia, natural sounds, natural environmental sounds, etc. In this embodiment, "voice control" means "control of the voice output from the speaker 21 of the voice output device 11." The camera 22 is exposed on the surface of the main body 50. The camera 22 captures an image of the surroundings of the voice output device 11, for example, the user's face, and outputs a signal corresponding to the captured image. The image captured by the camera 22 includes still images and videos.

[0016] The control unit 15 processes signals from the multiple pressure sensors 19, and stores the pressure information obtained by processing the signals from the pressure sensors 19 in the storage unit 16 in association with the second identification information. The control unit 15 also stores the user's voice collected by the microphone 20 in the storage unit 16 in association with the second identification information. The control unit 15 also stores the user's video captured by the camera 22 in association with the second identification information in the storage unit 16. The control unit 15 sends the pressure information, the user's voice, and the user's video to the user terminal 12 via the communication unit 17. The control unit 15 also causes the speaker 21 to output the output audio obtained from the user terminal 12.

[0017] (User terminal description) The user terminal 12 is a terminal operated and used by a user. The user terminal 12 includes a portable computer that the user can carry around, such as a smartphone or a tablet terminal. The user terminal 12 has a case main body 52 shown in FIG. 1. The case main body 52 is provided with a control unit 23, a memory unit 24, a communication unit 25, a display unit 26, an operation unit 27, a microphone 28, and a camera 29 shown in FIG. 3. The control unit 23 has an input port, an output port, an arithmetic processing circuit, a main memory circuit, etc.

[0018] The control unit 23 is connected to a memory unit 24, a communication unit 25, a display unit 26, an operation unit 27, a microphone 28, and a camera 29. The communication unit 25 communicates with the sound output tool 11, for example, by short-range waves. The communication unit 25 also communicates with the management server 13. The communication unit 25 is a device that receives and transmits radio waves, electricity, and light to communicate with other devices. The memory unit 24 stores non-transient applications, non-transient programs, information, data, etc. that are used for processing, such as control and judgment, performed by the control unit 23.

[0019] The display unit 26 has a configuration that can be seen by the user. Examples of the display unit 26 include a glass display, a liquid crystal display, a monitor, etc. The display unit 26 displays various screens, videos, images, text, icons, operation buttons, etc. The operation unit 27 is provided for manual operation by the user. When the user operates the operation unit 27, the control unit 23 loads applications stored in the memory unit 24 and executes various processes. The operation unit 27 includes switches, buttons, etc. physically provided on the case body 52 of the user terminal 12 shown in FIG. 1 as well as icons, operation buttons, etc. displayed on the display unit 26. The microphone 28 collects sounds uttered by the user and outputs a signal corresponding to the sounds. The camera 29 captures images of the surroundings of the audio output device 11, such as the user's face, and outputs a signal corresponding to the captured images. The images captured by the camera 29 include still images and videos.

[0020] The control unit 23 performs processing to send at least one of the pressure information and the user's voice information acquired from the voice output device 11 to the management server 13 via the communication unit 25. The control unit 23 also processes the signal from the microphone 28 and sends the user's voice information obtained by processing the signal from the microphone 28 to the management server 13 via the communication unit 25. The control unit 23 also performs processing to send the output voice information acquired from the management server 13 to the voice output device 11 via the communication unit 25. Note that a single voice output device 11 may be used by multiple different users.

[0021] (Management Server) The management server 13 is a computer managed by the operator of the voice control system 10. The management server 13 has a control unit 30, a memory unit 31, and a communication unit 32, as shown in FIG. 4. The communication unit 32 communicates with the user terminal 12. The communication unit 32 is a device that receives and transmits radio waves, electricity, and light to communicate with other devices. The memory unit 31 has a first information memory unit 40, a second information memory unit 41, a pressure information memory unit 42, an audio information memory unit 43, a video information memory unit 44, an output audio memory unit 45, and an application memory unit 46. The first information memory unit 40 stores data of first identification information.

[0022] The second information storage unit 41 stores data of second identification information. The pressure information storage unit 42 stores data of pressure information. The audio information storage unit 43 stores data of audio information. The video information storage unit 44 stores data of video information. The output audio storage unit 45 stores data of generated output audio. The application storage unit 46 stores non-transient applications and non-transient programs used in processes, such as control and judgment, performed by the control unit 30.

[0023] The learning data storage unit 47 stores learning data and the like used by the artificial intelligence unit 38. The learning data storage unit 47 also stores reference voice data for generating different output voices for each type of user voice information, reference voice data for generating different output voices for each type of user emotion information, reference voice data for generating different output voices for each type of touching the voice output tool 11, reference voice data for generating different output voices for each different time period when the user touches the voice output tool 11, reference voice data for generating different output voices for each type of voice output tool 11, reference voice data for generating different output voices for each user, and reference voice data for generating different output voices for each user. The emotion information storage unit 48 stores data on emotion information. The second identification information, pressure information, voice information, video information, emotion information, and output voices will be described later.

[0024] The control unit 30 is connected to the memory unit 31 and the communication unit 32. The control unit 30 has an input port, an output port, an arithmetic processing circuit, a main memory circuit, etc. The control unit 30 functions as a first identification unit 33, a second identification unit 34, a pressure information processing unit 35, a sound processing unit 36, a video processing unit 37, and an artificial intelligence unit 38 by reading and starting non-transient applications and non-transient programs stored in the memory unit 31. The first identification unit 33 individually identifies users who use the sound output tool 11, i.e., who touch it with their hands, by using first identification information, and stores the first identification information in the memory unit.

[0025] The first identification information includes, for example, a user ID, a user name, the user's age, the user's gender, etc., assigned to each user who owns or touches the sound output tool 11. The second identification unit 34 identifies each sound output tool 11 individually using the second identification information assigned to each sound output tool 11, and stores the second identification information in the storage unit 31. The second identification information includes an identification ID assigned to each sound output tool 11, the type of sound output tool 11, etc. The second identification information is associated with the first identification information of the user who owns the sound output tool 11.

[0026] The pressure information processing unit 35 processes the pressure information acquired from the user terminal 12 and stores it in the storage unit 31. The audio processing unit 36 processes the user's voice acquired from the user terminal 12 and stores it in the storage unit 31. The video processing unit 37 processes the user's video information acquired from the user terminal 12, and stores the user's video information in the storage unit 31 in association with the first identification information and the second identification information.

[0027] The artificial intelligence unit 38 creates history data and update data for the various information stored in the memory unit 31 based on the various information acquired by the control unit 30, the various information stored in the memory unit 31, and the learning data, and stores the created data in the memory unit 31. The artificial intelligence unit 38 also generates an output sound based on the pressure information, the user's voice information, the user's video information, and the learning data. The output sound is a sound to be output from the speaker 21 of the sound output tool 11. The artificial intelligence unit 38 stores the generated output sound in the memory unit 31. The output sound stored in the memory unit 31 is associated with the first identification information and the second identification information.

[0028] The artificial intelligence unit 38 is an artificial intelligence (AI) that includes transformers such as GPT (Generative Pre-trained Transformer) and BERT (Bidirectional Encoder Representations from Transformers), and language models such as recurrent neural networks, and is capable of performing processing as generative artificial intelligence.

[0029] The language model is an example of a learning model based on a machine learning algorithm. Specific examples of machine learning algorithms include nearest neighbor methods, naive Bayes methods, decision trees, support vector machines, and deep learning using neural networks. The artificial intelligence unit 38 can apply the above algorithms as appropriate.

[0030] The artificial intelligence unit 38 may have a trained model constructed by a learning method such as supervised learning, unsupervised learning, or self-supervised learning. In supervised learning, machine learning is performed using training data including training data. The training data is composed of pairs of input data for learning and output data including correct answer data. In addition, the language model includes a model trained for a specific task and a general-purpose model that can be used for a wide range of tasks.

[0031] (An example of voice control method) FIG. 5 is a flowchart showing an example of a voice control method performed in the voice control system 10.

[0032] (Audio output device processing) The control unit 15 of the sound output device 11 reads applications stored in the storage unit 16 and executes various processes. For example, in step S10, the control unit 15 acquires signals from a plurality of pressure sensors 19, a signal from the microphone 20, and a signal from the camera 22. The control unit 11 associates the signals from the pressure sensors 19, the signals from the microphone 20, and the signals from the camera 22 acquired in step S10 with second identification information, and transmits them to the user terminal 12 in step S11.

[0033] In step S12 following step S11, the control unit 15 of the audio output device 11 acquires audio to be output from the user terminal. In step S13, the control unit 15 of the audio output device 11 causes the speaker 21 to output the audio to be output.

[0034] (User terminal processing) In step S20, the user operates the user terminal 12 and inputs the first identification information and the second identification information. In step S20, the user terminal 12 acquires the first identification information and the second identification information and stores the acquired first identification information and second identification information in the storage unit 24. In step S21 following step S20, the user terminal transmits the first identification information and the second identification information to the management server 13.

[0035] In step S22 following step S21, the user terminal 12 acquires a signal from the pressure sensor 19, a signal from the microphone 20, and a signal from the camera 22 from the sound output tool 11. The user terminal can store the signal from the pressure sensor 19, the signal from the microphone 20, and the signal from the camera 22 acquired from the sound output tool 11 in the memory unit 24 in association with the first identification information and the second identification information.

[0036] Furthermore, if the user terminal 12 is equipped with a camera 29, the user terminal 12 can acquire a signal from the camera 29 in step S22 and store the signal in the storage unit 24. If the user terminal 12 is equipped with a microphone 28, the user terminal 12 can acquire a signal from the microphone 28 in step S22 and store the signal in the storage unit 24. The user terminal 12 can store the signal from the camera 29 and the signal from the microphone 28 in the storage unit 24 in association with the first identification information and the second identification information.

[0037] In step S23 following step S22, the user terminal 12 associates the signal from the pressure sensor 19, the signal from the microphone 20, and the signal from the camera 22 with the first identification information and the second identification information and transmits them to the management server 13. In addition, in step S23, the user terminal 12 can also associate the signal from the microphone 28 and the signal from the camera 29 with the first identification information and the second identification information and transmit them to the management server 13.

[0038] In step S24 following step S23, the user terminal 12 acquires the audio to be output from the management server 13. The audio to be output acquired by the user terminal 12 is associated with the first identification information and the second identification information. In step S25 following step S24, the user terminal 12 transmits the audio to be output to the audio output tool 11 identified by the second identification information.

[0039] (Management server processing) In step S30, the management server 13 acquires the first identification information and the second identification information from the user terminal 12. In step S30, the management server 13 also stores the first identification information and the second identification information in the storage unit 31. In step S31 following step S30, the management server 13 acquires signals from the pressure sensor 19, the microphones 20 and 28, and the cameras 22 and 29 from the user terminal 12 in association with the first identification information and the second identification information.

[0040] In step S32, the control unit 30 of the management server 13 processes the various signals, first identification information, and second identification information acquired from the user terminal 12, which are stored in the storage unit 31, and stores the processing results in the storage unit 31. The control unit 30 of the management server 13 processes the signals of the pressure sensor 19 and the signals of the cameras 22 and 29 in step S32, thereby acquiring pressure information such as that shown in Fig. 6. The pressure information includes the date and time period when the user touched the sound output device 11, the number of times the user touched the sound output device 11 within a predetermined time period, the location on the main body 50 touched by the user's hand, the pressures applied to each of the multiple pressure sensors 19, and the amount of change in pressure, etc.

[0041] The control unit 30 of the management server 13 processes the signals from the microphones 20 and 28 in step S32 to acquire user voice information such as that shown in FIG. 6. The user voice information includes the date and time of the user's interaction while touching the voice output device 11, the number of times the user interacted with the voice output device 11 within a predetermined time period, the content of the user's interaction, etc. The interaction includes words spoken by the user alone, conversations and casual conversations between the user and the voice output device 11, etc. The control unit 30 of the management server 13 processes the signals from the cameras 22 and 29 in step S32 to acquire user video information such as that shown in FIG. 7. The video information includes the video, the date and time of capture, and the time period of capture. In addition, in step S32, the control unit 30 of the management server 13 stores the pressure information, the user's voice information, and the user's video information in the storage unit 31 in association with the user's identification information and second identification information.

[0042] In step S32, the control unit 30 of the management server 13 determines how the audio output device 11 is being touched based on the amount of change in pressure. FIG. 8 shows the relationship between the amount of change in pressure applied to the audio output device 11 and how the audio output device 11 is being touched, which is an example of how the audio output device 11 is being used, assuming that the pressure applied to the audio output device 11 is equal to or less than a predetermined value. The amount of change in pressure applied to the audio output device 11 is determined by the "pattern of change in pressure value." Furthermore, the "way of touching" the audio output device 11 is determined based on the amount of change in pressure. For example, when the pattern of change in pressure value is such that "the threshold is set to a value where the amount of change in pressure in the superficial layer is small, and the threshold is set to a value where the amount of change in pressure in the deep layer is considered to be zero," the determination is made as "the amount of change in pressure in the superficial layer is small, and there is no change in pressure in the deep layer." Then, the control unit 30 can determine that "the audio output device 11 has been stroked" when "the amount of change in pressure in the superficial layer is small, and there is no change in pressure in the deep layer."

[0043] Furthermore, when the pattern of pressure value change is such that "the threshold is set to a value where the amount of change in pressure in the superficial layer is large, and the threshold is set to a value where the amount of change in pressure in the deep layer is considered to be zero," it is determined that "the amount of change in pressure in the superficial layer is large, and the amount of change in pressure in the deep layer is zero." Then, the control unit 30 can determine that "the voice output instrument 11 has been pinched" when "the amount of change in pressure in the superficial layer is large, and the amount of change in pressure in the deep layer is zero."

[0044] Furthermore, when the pattern of pressure value change is such that "the threshold is set to a value where the amount of change in pressure in the superficial layer is large, and the threshold is set to a value where the amount of change in pressure in the deep layer is large," it is determined that "the amount of change in pressure in the superficial layer is large, and the amount of change in pressure in the deep layer is large." Then, the control unit 30 can determine that "the sound output tool 11 has been kneaded" when "the amount of change in pressure in the superficial layer is large, and the amount of change in pressure in the deep layer is large."

[0045] 8, the control unit 30 can also determine that "the user is sleeping using the sound output device 11 as a body pillow" when all of the multiple pressure sensors 19 receive pressure, and the pressure received at the surface layer is higher than a predetermined value, and the pressure received at the deep layer is higher than a predetermined value. The control unit 30 stores the manner of touching the sound output device 11 in the storage unit 31 in association with the first identification information and the second identification information.

[0046] Furthermore, in step S32, the control unit 30 can estimate the emotional information (psychological information) of the user using the voice output device 11 by having the generative artificial intelligence perform processing, inference, judgment, etc. using machine learning using the user's video information, the user's voice information, and information, data, etc. stored in the memory unit 31.

[0047] As shown in FIG. 7 , the user's emotional information includes satisfaction, relaxation, calmness, surprise, fatigue, liveliness, activity, excitement, and enjoyment. The user's emotional information also includes satisfaction, relaxation level, level of calmness, level of surprise, level of fatigue, level of liveliness, level of activity, level of excitement, and level of enjoyment. This emotional information can be determined and acquired by processing the eyelid opening, gaze, eyebrow movement, presence or absence of nose wrinkles, mouth movement, mouth opening, pupil opening, and the like, contained in the video of the user's upper body including the face. These emotions can also be estimated from the user's upper body movements, such as whether or not the head is tilted, whether or not the user nods, and hand movements. The estimated emotions are each quantified, and a table or graph can be generated.

[0048] The technology for determining and acquiring a user's emotional information based on video information including the user's face is a publicly known technology, as shown in, for example, Patent Publication No. 6868422, Patent Publication No. 6703893, Patent Publication No. 6042015, Patent Publication No. 3953024, Patent Publication No. 4458888, JP 2015-229040, JP 2018-32164, JP 2022-139436, JP 2020-184216, etc., and therefore a detailed description thereof will be omitted.

[0049] Furthermore, the technology for determining and acquiring a user's emotional information from the user's voice information is a publicly known technology, as shown in, for example, Patent Publication No. 2874858, Patent Publication No. 3676969, Patent Publication No. 3676981, Patent Publication No. 4580190, Patent Publication No. 4670431, JP 2012-59107, JP 2015-141428, JP 2021-11071, etc., and therefore a detailed explanation thereof will be omitted.

[0050] In step S33, the control unit 30 of the management server 13 starts up the programs and applications stored in the storage unit 31, and generates audio to be output based on the processing results of step S32 and various data stored in the storage unit 31. In other words, the control unit 30 generates audio to be output by processing pressure information previously stored in the storage unit 31, user voice information previously stored in the storage unit 31, user video information previously stored in the storage unit 31, as well as newly acquired pressure information, newly acquired user voice information, new user video information, etc.

[0051] Specifically, the control unit 30 of the management server 13 generates an output sound based on pressure information, user voice information, user video information, emotional information, etc., using prompts such as the dialogue between the user and the speaker 21 of the sound output tool 11 and the state of the user touching the sound output tool 11 as prompts, and the generative artificial intelligence responds to the prompts. Note that the various pieces of information are associated with first identification information and second identification information. Furthermore, the generated output sound is associated with the first identification information and second identification information.

[0052] For example, the output audio generated for a younger user includes content encouraging the user to be active, while the output audio generated for an older user includes content expressing appreciation and gratitude to the user. Furthermore, the output audio generated when the audio output tool 11 is stroked includes content expressing gratitude to the user and that it feels good to be stroked. Furthermore, the output audio generated when the audio output tool 11 is pinched includes content expressing a desire for the user to stop and content expressing a feeling of pain. Furthermore, the output audio generated when the audio output tool 11 is kneaded includes content expressing gratitude to the user.

[0053] Furthermore, the output voice generated when the time period when the voice output device 11 is touched is before 9:00 a.m. includes a message encouraging the user to have a good day. Furthermore, the output voice generated when the time period when the voice output device 11 is touched is after 6:00 p.m. includes a message praising the user's actions for the day.

[0054] Furthermore, the output sound generated when the user who touches the sound output tool 11 feels tired includes content that praises the user. The output sound generated when the user who touches the sound output tool 11 feels satisfied includes content that praises the user. Furthermore, the output sound generated when the user is sleeping with the sound output tool 11 as a pillow includes music or natural sounds that relax the user.

[0055] Furthermore, the control unit 30 determines the content of the voice to be generated based on information such as the user's age, sex, the period of time since the user started using the voice output device 11, the time period when the user came into contact with the voice output device 11, and the type of voice output device 11. The control unit 30 can determine the level of satisfaction from the user's emotional information and determine the content of the voice to be generated so that the user's level of satisfaction is relatively high.

[0056] Furthermore, the control unit 30 can vary the content of the sound to be generated depending on the time of day when the user is touching the sound output device 11. For example, when the user is touching the sound output device 11 in the morning, for example, between 6:00 AM and 10:00 AM, the control unit 30 can determine the content of the sound to be generated so that the user's emotions become relatively lively, active, and excited. This makes it possible to obtain a situation in which the user's behavior is likely to become active. Furthermore, when the user is touching the sound output device 11 in the afternoon, for example, between 9:00 PM and 12:00 PM, the control unit 30 can determine the content of the sound to be generated so that the user's level of relaxation increases. This makes it possible to obtain a situation in which the user is likely to fall asleep.

[0057] Furthermore, the control unit 30 can vary the content of the sound to be output depending on the age group of the user touching the sound output device 11. For example, the control unit 30 can determine the content of the sound to be output so that the level of enjoyment increases when the user touching the sound output device 11 belongs to a relatively young age group, for example, 10 years old or younger. Furthermore, the control unit 30 can determine the content of the sound to be output so that the level of satisfaction and calmness increases when the user touching the sound output device 11 belongs to a relatively older age group, for example, 20 years old or older.

[0058] The output voice is updated every time the control unit 30 acquires new information, and is stored together with past history in the storage unit 31. In step S34 following step S33, the control unit 30 of the management server 13 transmits, from the generated output voices, for example, the latest output voice to the user terminal 12.

[0059] (Effects of this embodiment) The control unit 30 of the management server 13 executes a process of acquiring voice information uttered by the user in step S31. Furthermore, the control unit 30 of the management server 13 executes a process of determining the generation of an output voice to be output from the speaker 21 in accordance with the user's contact state with the voice output tool 11 and the voice information uttered by the user in step S33. When the output voice is output from the speaker 21 of the voice output tool 11 in step S13, the user's emotional information is determined from the video information of the cameras 22, 29 and from the voice information uttered by the user after step S13.

[0060] Furthermore, the control unit 15 of the voice output device 11 can repeatedly execute the processes of steps S10, S11, S12, and S13 shown in Fig. 5. Furthermore, the control unit 23 of the user terminal 12 can repeatedly execute the processes of steps S22, S23, S24, and S25 shown in Fig. 5. Furthermore, the control unit 30 of the management server 13 can repeatedly execute the processes of steps S31, S32, S33, and S34 shown in Fig. 5. In other words, a conversation or chat is established between the user and the voice output device 11.

[0061] In this way, the voice control system 10 can adjust the content of the voice output from the speaker 21 of the voice output device 11 to correspond to the pressure state of the user's touch on the voice output device 11 and the voice uttered by the user. Therefore, the content of the voice to be output from the speaker 21 can be adjusted to match the user's touch state and the voice information uttered by the user, which increases the user's sense of satisfaction and prevents the user from feeling uncomfortable. Furthermore, the user can have a dialogue with the voice output device 11, which gives the user a sense of fulfillment.

[0062] Furthermore, the control unit 30 of the management server 13 further executes a third process in step S31 to acquire emotional information of the user. Therefore, the second process executed by the control unit 30 in step S32 includes a process of determining the content of the output sound to be output from the speaker 21 in accordance with the user's contact state, the voice information emitted by the user, and the user's emotional information. In other words, by performing generative artificial intelligence processing, the control unit 30 can generate output sound of a type and characteristics that can relatively increase the user's satisfaction. Therefore, the content of the output sound to be output from the speaker 21 of the voice output device 11 can be made appropriate for the user's contact state, the voice information emitted by the user, and the user's emotional information, thereby increasing the user's satisfaction and preventing the user from feeling uncomfortable.

[0063] Furthermore, the artificial intelligence unit 38 of the control unit 30 generates output voice by performing generative artificial intelligence processing based on various information and data. Therefore, even if the amount of data and information pre-stored in the memory unit 31 is insufficient, the control unit 30 can generate original output voice. Therefore, the management server 13 can suppress an increase in the amount of data and information pre-stored in the memory unit 31, and can suppress an increase in the capacity of the memory unit 31. Furthermore, the management server 13 can suppress an increase in the amount of data and information handled by the arithmetic processing circuit and main memory circuit of the control unit 30. Therefore, the control unit 30 can relatively increase the processing speed for generating output voice.

[0064] Furthermore, the artificial intelligence unit 38 of the control unit 30 performs generative artificial intelligence processing based on the state of contact between the user and the sound output device 11, the user's voice information, and the user's emotional information, to generate a sound for output. Therefore, the sound output device 11 can increase user satisfaction. Also, the control unit 30 can increase the accuracy of generating a sound for output to increase user satisfaction. Furthermore, the sound output device 11 can increase the accuracy of outputting a sound for output to increase user satisfaction.

[0065] (supplementary explanation) The voice control system is not limited to the configuration and functions disclosed in this embodiment, and various modifications are possible without departing from the spirit thereof. For example, the voice output device 11 may include a display unit 26 and an operation unit 27. Furthermore, the control unit 15 of the voice output device 11 may have the configuration of the control unit 30. The storage unit 16 of the voice output device 11 may have the configuration of the storage unit 31. When the voice output device 11 is configured in this manner, the voice output device 11 can perform the processing of step S20 even when the voice output device 11 is not connected to the user terminal 12. Furthermore, the voice output device 11 can perform the processing of steps S31, S32, and S33.

[0066] Furthermore, the control unit 23 of the user terminal 12 may have the configuration of the control unit 30. The storage unit 24 of the user terminal 12 may have the configuration of the storage unit 31. When the user terminal 12 is configured in this manner, the processes of steps S32 and S33 can be performed in the user terminal 12 even when the user terminal 12 is not connected to the management server 13.

[0067] 5 may be stored in any one of the storage units 16, 24, and 31 depending on the configuration of the voice control system 10. The non-transient application includes a non-transient program.

[0068] Furthermore, a storage medium 51 that can be attached to and detached from the main body 50 of the audio output instrument 11 may be provided. Furthermore, a storage medium 53 that can be attached to and detached from the case main body 52 of the user terminal 12 may be provided. These storage media 51, 53 include DVDs (Digital Versatile Discs), CDs (Compact Discs), flash memories, etc. A non-transitory application and a non-transitory program that execute the audio output control method shown in FIG. 5 are stored in the storage media 51, 53.

[0069] The user terminal 12 may be a portable computer or a stationary computer, such as a desktop computer. The management server 13 may be a single computer or a distributed computer made up of multiple devices. The computer connected to the user terminal 12 via the network 14 is not limited to the management server 13, but may also be a supercomputer, a mainframe, a workstation, or the like.

[0070] An example of the technical meaning of the matters described in this embodiment is as follows: The audio output device 11 is an example of an audio output device. The user terminal 12 and the management server 13 are each an example of an audio control device. The main body 50 is an example of a main body. The speaker 21 is an example of a speaker 21. The control units 15, 23, and 30 are examples of a control unit. Step S32 is an example of a first process, a second process, and a fourth process. Step S33 is an example of a third process. The pressure sensor 19 is an example of a pressure sensor.

[0071] In addition, this embodiment discloses a recording medium storing an application that causes a control unit provided in an audio output device to execute a process for determining the output audio to be output from a speaker provided in the audio output device, wherein the application causes the control unit to execute a first process for acquiring audio information uttered by a user using the audio output device, a second process for acquiring the user's contact state with the audio output device, and a third process for determining the content of the output audio to be output from the speaker based on the user's contact state with the audio output device and the audio information uttered by the user.

[0072] The various control units described in this embodiment can also be defined as control circuits, controllers, and control means. The various storage units can also be defined as storage circuits, storage units, and storage means. The communication units can also be defined as communication circuits, communication units, and communication means. Furthermore, the various identification units can also be defined as identification circuits, identifiers, and identification means. Furthermore, the various processing units can also be defined as processing circuits, processors, and processing means. Furthermore, the artificial intelligence unit can also be defined as an artificial intelligence circuit. [Industrial Applicability]

[0073] The present embodiment can be used as a sound output tool, a sound control device, a sound control method, and an application that can output sound according to the state of contact by the user. [Explanation of symbols]

[0074] 11...sound output device, 12...user terminal, 13...management server, 15, 23, 30...control unit, 19...pressure sensor, 21...speaker, 50...main body

Claims

1. The main body and a speaker provided on the main body and configured to output audio to a user; a pressure sensor provided on the main body and configured to detect a pressure applied by the user to the main body; a control unit provided in the main body and configured to determine an output sound to be output from the speaker; An audio output tool having: The control unit a first process for acquiring voice information uttered by the user; a second process of acquiring a contact state of the user with the main body determined from the pressure detected by the pressure sensor; a third process for determining the content of the sound to be output from the speaker in accordance with the state of contact of the user with the main body and the sound information uttered by the user; An audio output tool that performs the above.

2. 2. The audio output tool according to claim 1, the control unit further executes a fourth process of acquiring emotion information of the user; The third processing executed by the control unit includes processing for determining the content of the output audio to be output from the speaker in accordance with the contact state of the user, audio information emitted by the user, and emotional information of the user.

3. 3. The audio output tool according to claim 2, The third processing executed by the control unit includes processing for changing the content of the output audio output from the speaker to content that relatively increases the satisfaction level included in the user's emotional information.

4. 3. The audio output tool according to claim 2, The third process executed by the control unit includes a process of varying the content of the output sound output from the speaker depending on the age group of the user.

5. 3. The audio output tool according to claim 2, The third process executed by the control unit includes a process of varying the content of the output sound output from the speaker depending on the time of day when the user touches the main body.

6. A voice control device connected via a network to a voice output tool that a user touches, and having a control unit that determines a voice to be output from a speaker provided in the voice output tool, The control unit a first process for acquiring voice information uttered by the user; a second process of acquiring a contact state of the user with respect to the voice output tool; a third process for determining the content of the sound to be output from the speaker in accordance with the state of contact of the user with the sound output tool and sound information uttered by the user; A voice control device that performs the above.

7. a speaker provided on the main body and configured to output audio to a user; a pressure sensor provided on the main body and configured to detect a pressure applied by the user to the main body; a control unit provided in the main body and configured to determine an output sound to be output from the speaker; an audio output tool having The control unit a first process for acquiring voice information uttered by the user; a second process of acquiring a contact state of the user with the main body determined from the pressure detected by the pressure sensor; a third process for determining the content of the sound to be output from the speaker in accordance with the state of contact of the user with the main body and the sound information uttered by the user; Implementing a voice control method.

8. a voice control device connected via a network to a voice output tool that a user touches, the voice control device having a control unit that determines a voice to be output from a speaker provided in the voice output tool; The control unit a first process for acquiring voice information uttered by the user; a second process of acquiring a contact state of the user with respect to the voice output tool; a third process for determining the content of the sound to be output from the speaker in accordance with the state of contact of the user with the sound output tool and sound information uttered by the user; Implementing a voice control method.

9. An application that causes a control unit provided in a sound output tool to execute a process for determining an output sound to be output from a speaker provided in the sound output tool, The control unit a first process of acquiring voice information uttered by a user using the voice output tool; a second process of acquiring a contact state of the user with respect to the voice output tool; a third process for determining the content of the sound to be output from the speaker based on the state of contact of the user with the sound output tool and sound information uttered by the user; Run the application.

Citation Information

Patent Citations

  • Information processing device, information processing method, and program

    JP2019008570A

  • Method, system and computer program product for wearable computing device audio interface (wearable computing device audio interface)

    JP2022054447A

  • Display device using magnetic fluid

    JP7360756B1

  • Information processing device, information processing method, and program

    WO2019087495A1

  • Content provision system, content provision method, and storage medium

    WO2021157380A1

Cited By

  • Information processing device, information processing method, program, and storage medium

    JP7919790B1