Audio processing device, audio processing system, audio processing method, and audio processing program

The voice processing system addresses the challenge of selecting desired voices by separating and emphasizing them based on user motion, allowing easy and intuitive voice selection and listening.

WO2025142347A1PCT designated stage expired Publication Date: 2025-07-03KYOCERA CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/042764
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-25
Filing Date
2024-12-03
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Existing voice processing systems struggle to effectively emphasize voices that a user wants to hear amidst a mixture of ambient sounds, making it difficult for users to easily select and listen to desired voices.

Method used

A voice processing system that includes a sound collection device, a detection device, and a voice output device, which separates and emphasizes voices based on the user's selected direction and distance using motion data to generate reproduction data with weighted emphasis on desired voices.

Benefits of technology

Enables users to intuitively select and easily hear desired voices by determining the selection direction and distance based on user movement, enhancing the user's listening experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024042764_03072025_PF_FP_ABST
    Figure JP2024042764_03072025_PF_FP_ABST
Patent Text Reader

Abstract

This audio processing device comprises: an acquisition unit that acquires audio data obtained by collecting sound occurring around a user, and motion data obtained by detecting the user's movement; and a control unit that determines a selection direction on the basis of the motion data, and generates reproduction data by giving greater weight to sound occurring in the selection direction than to sound occurring in other directions.
Need to check novelty before this filing date? Find Prior Art

Description

Audio processing device, audio processing system, audio processing method, and audio processing program Cross-reference to related applications

[0001] This application claims priority from Japanese Patent Application No. 2023-218537 (filed December 25, 2023), the entire disclosure of which is incorporated herein by reference.

[0002] The present disclosure relates to a voice processing device, a voice processing system, a voice processing method, and a voice processing program.

[0003] As described in Patent Document 1, there is known a device that collects surrounding sounds with a microphone, records the recorded sounds, and plays back the recorded sounds in response to a user operation.

[0004] International Publication No. 2023 / 140149

[0005] According to an embodiment of the present disclosure, there is provided a sound processing device including: an acquisition unit; and a control unit. The acquisition unit acquires sound data obtained by collecting sound generated around a user and motion data obtained by detecting the user's movement. The control unit determines a selection direction to be selected by the user based on the motion data, and generates playback data by assigning a greater weight to sound generated in the selection direction among the sound data than to sound generated in directions other than the selection direction.

[0006] An audio processing system according to an embodiment of the present disclosure includes a sound collection device, a detection device, an audio processing device, and an audio output device. The sound collection device collects audio generated around a user and generates audio data. The detection device detects the user's movements and outputs the motion data. The sound processing device acquires the audio data from the sound collection device and acquires the motion data from the detection device. The audio processing device determines a selection direction to be selected by the user based on the motion data. The audio processing device generates playback data by assigning a greater weight to audio generated in the selection direction among the audio data than to audio generated in directions other than the selection direction. The audio output device outputs the playback data to the user as audio.

[0007] An audio processing method according to an embodiment of the present disclosure includes an audio processing device acquiring audio data obtained by collecting audio generated around a user and motion data obtained by detecting motion of the user. The audio processing method also includes the audio processing device determining a selection direction to be selected by the user based on the motion data. The audio processing method also includes the audio processing device generating playback data by assigning a greater weight to audio generated in the selection direction of the audio data than to audio generated in directions other than the selection direction.

[0008] An audio processing program according to an embodiment of the present disclosure causes a processor to acquire audio data obtained by collecting audio occurring around a user and motion data obtained by detecting the user's movements. The audio processing program also causes the processor to determine a selection direction to be selected by the user based on the motion data. The audio processing program also causes the processor to generate playback data by assigning a greater weight to audio occurring in the selection direction among the audio data than to audio occurring in directions other than the selection direction.

[0009] Fig. 1 is a schematic diagram showing a configuration example of a speech processing system according to the present disclosure; Fig. 2 is a block diagram showing a configuration example of a speech processing system according to the present disclosure; Fig. 3 is a flowchart showing an example procedure of a speech processing method; Fig. 4 is a flowchart showing an example procedure for setting a weight in consideration of distance; Fig. 5 is a flowchart showing an example procedure for registering a reference point; Fig. 6 is a flowchart showing an example procedure for determining a selection direction;

[0010] The sounds occurring around the user include sounds that the user wants to hear and sounds that the user does not need to hear. It is required that the sounds that the user wants to hear can be emphasized. Considering user convenience, it is required that the user can easily select the sounds that they want to hear. According to the sound processing device, sound processing system, sound processing method, and sound processing program according to the present disclosure, the sound that the user wants to hear can be easily selected.

[0011] (Configuration example of audio processing system 1) As illustrated in Figs. 1 and 2, an audio processing system 1 according to one embodiment includes an audio processing device 10, a sound collection device 20, a detection device 30, and an audio output device 40. In this embodiment, the audio processing device 10, the sound collection device 20, and the audio output device 40 are configured as an integrated unit and are worn on the user's ear. The detection device 30 is worn on the user's arm. As will be described later, the combination of the devices may be changed as appropriate. Furthermore, the devices may be worn on other parts of the user's body.

[0012] In the sound processing system 1, the sound collection device 20 detects sounds occurring around the user, and the detection device 30 detects the movement or orientation of the user's arm.

[0013] When the sound generated around the user is a mixture of independent sounds generated from a plurality of sound sources, the sound processing device 10 separates the independent sounds from the sound detected by the sound collection device 20 and associates each independent sound with the direction in which the sound source of the independent sound is located.

[0014] The voice processing device 10 identifies a direction selected by the user based on the detection result of the user's arm movement, and generates playback data that emphasizes an independent voice associated with the direction selected by the user. The direction selected by the user is also referred to as a selection direction.

[0015] The audio output device 40 outputs the reproduced data to the user as audio.

[0016] As described above, by operating the voice processing system 1, the user can select and listen to an independent voice that is generated in a direction indicated by the movement of the user's arm from a mixture of multiple independent voices that are generated around the user.

[0017] An example of the configuration of the voice processing system 1 will be described below.

[0018] <Speech Processing Device 10 > The speech processing device 10 includes a control unit 11 , a storage unit 12 , and an interface 13 .

[0019] The control unit 11 may be configured to include at least one processor to perform various functions, such as the functions described above. The processor may include, for example, a CPU (Central Processing Unit) or a DSP (Digital Signal Processor). The processor may include a general-purpose processor that loads a specific program to perform a specific function, or a dedicated processor specialized for a specific process. The dedicated processor may include an ASIC (Application Specific Integrated Circuit). The processor may include a PLD (Programmable Logic Device). The PLD may include an FPGA (Field-Programmable Gate Array). The control unit 11 may be either a SoC (System-on-a-Chip) or a SiP (System-in-a-Package) in which one or more processors work together.

[0020] The processor may be implemented as a single integrated circuit, also referred to as an IC (Integrated Circuit). The processor may be implemented as multiple communicatively connected integrated circuits and discrete circuits. The processor may also be implemented based on various other known technologies.

[0021] The control unit 11 may be configured to include at least one of software and hardware resources, for example, or may be configured by specific means in which software and hardware resources work together.

[0022] The storage unit 12 stores various information or programs used in the voice processing device 10. The storage unit 12 may store, for example, a program executed by the control unit 11 or the results of processing executed by the control unit 11. The storage unit 12 may function as a work memory for the control unit 11.

[0023] The storage unit 12 may include, for example, at least one of a RAM (Random Access Memory) and a ROM (Read Only Memory). The storage unit 12 may be configured to include, for example, a semiconductor memory or the like, but is not limited thereto, and may be configured to include any storage device. The storage unit 12 may be configured as, for example, a storage medium such as a memory card inserted into the audio processing device 10. The storage unit 12 may be an internal memory of the CPU used as the control unit 11, or may be a memory connected to the control unit 11 as a separate unit.

[0024] The interface 13 is configured to allow the voice processing device 10 to accept input of information and output information. When the interface 13 accepts input of information and acquires information, it is also referred to as an acquisition unit. When the interface 13 outputs information, it is also referred to as an output unit. The interface 13 may be configured to enable the voice processing device 10 to communicate with other devices via a network or without a network. The interface 13 may include a communication device or a communication module. The interface 13 may be configured to be connectable to an external communication device or a communication module. The communication device or the communication module may be compatible with mobile communication standards such as 4G (4th Generation) or 5G (5th Generation). The communication device or the communication module may be compatible with communication standards such as LAN (Local Area Network). The communication device or the communication module may be compatible with wired or wireless communication standards. The communication device or the communication module is not limited to these and may be compatible with various communication standards.

[0025] The interface 13 may be configured to include an input device that accepts input of operations, information, data, etc. from a user. The interface 13 may be configured to be connectable to an external input device. The input device may be configured to include, for example, a touch panel or touch sensor, or a pointing device such as a mouse. The input device may be configured to include physical keys.

[0026] The interface 13 may include an output device that outputs information, data, or the like to the user. The interface 13 may be configured to be connectable to an external output device. The output device may include, for example, a display device that outputs visual information such as images, characters, or graphics. The display device may include, for example, a liquid crystal display (LCD), an organic electroluminescence (EL) display, an inorganic electroluminescence (EL) display, or a plasma display panel (PDP). The display device is not limited to these displays and may include various other display types. The display device may include a light-emitting device such as an LED (light-emitting diode) or an LD (laser diode). The display device may include various other devices. The display device may be a touch panel display that inputs information by touching it with a finger or the like. The output device may include, for example, a vibration device such as a vibrator that outputs tactile information such as vibration. The output device is not limited to these examples and may include various other devices.

[0027] <Sound collection device 20> The sound collection device 20 collects sounds generated around the sound collection device 20 and outputs sound data. The sound collection device 20 includes a microphone 21, a control unit 22, and an interface 23.

[0028] The microphone 21 detects sounds occurring around the sound collection device 20. The number of microphones 21 may be one or two or more. The microphone 21 may be configured to have directionality. When the number of microphones 21 is two or more, each of the two or more microphones 21 may detect sounds occurring around the device and output the sound detection results. The microphone 21 may be configured as a microphone array.

[0029] The control unit 22 generates audio data based on the audio detected by the microphone 21 and outputs the audio data from the interface 23. The control unit 22 may be configured to include a processor or a dedicated circuit. The control unit 22 may be configured in the same manner as or similar to the control unit 11 of the audio processing device 10.

[0030] The sound collection device 20 may include a storage unit that stores various information or programs used by the control unit 22 of the sound collection device 20. The storage unit of the sound collection device 20 may store, for example, a program executed by the control unit 22 or the results of processing executed by the control unit 22. The storage unit of the sound collection device 20 may function as a work memory for the control unit 22. The storage unit of the sound collection device 20 may be configured in the same manner as or similar to the storage unit 12 of the sound processing device 10.

[0031] The interface 23 includes a communication device or a communication module configured to be able to communicate with the interface 13 of the audio processing device 10. The interface 23 may be configured the same as or similar to the interface 13 of the audio processing device 10.

[0032] <Detection Device 30 > The detection device 30 includes a motion sensor 31 , a control unit 32 , and an interface 33 .

[0033] The motion sensor 31 may include an angular velocity sensor or a gyro sensor. The motion sensor 31 may include an acceleration sensor. The motion sensor 31 may include an IMU (Inertial Measurement Unit) that detects translational motion and rotational motion in three mutually orthogonal axial directions, i.e., three-dimensional inertial motion. The motion sensor 31 is attached to the user's body and detects the movement of the part where the motion sensor 31 is attached. When the motion sensor 31 is attached to the user's arm, for example, the motion sensor 31 detects the movement of the user's arm.

[0034] The control unit 32 acquires the detection results of the movement of the wearing part, generates data on the user's movement, and outputs it from the interface 33. The control unit 32 may be configured to include a processor or a dedicated circuit. The control unit 32 may be configured identically or similarly to the control unit 11 of the sound processing device 10.

[0035] The detection device 30 may include a storage unit that stores various information or programs used by the control unit 32 of the detection device 30. The storage unit of the detection device 30 may store, for example, a program executed in the control unit 32 or the results of processing executed in the control unit 32. The storage unit of the detection device 30 may function as a work memory for the control unit 32. The storage unit of the detection device 30 may be configured in the same manner as or similar to the storage unit 12 of the voice processing device 10.

[0036] The interface 33 includes a communication device or a communication module configured to be able to communicate with the interface 13 of the audio processing device 10. The interface 33 may be configured identically to or similarly to the interface 13 of the audio processing device 10.

[0037] <Audio Output Device 40 > The audio output device 40 includes a speaker 41 , a control unit 42 , and an interface 43 .

[0038] The speaker 41 outputs sound and may be configured as an earphone or a headphone.

[0039] The control unit 42 generates an audio signal based on the playback data generated by the audio processing device 10 and outputs the audio from the speaker 41. The control unit 42 may be configured to include a processor or a dedicated circuit. The control unit 42 may be configured identically or similarly to the control unit 11 of the audio processing device 10.

[0040] The audio output device 40 may include a storage unit that stores various information or programs used by the control unit 42 of the audio output device 40. The storage unit of the audio output device 40 may store, for example, a program executed by the control unit 42 or the results of processing executed by the control unit 42. The storage unit of the audio output device 40 may function as a work memory for the control unit 42. The storage unit of the audio output device 40 may be configured in the same manner as or similar to the storage unit 12 of the audio processing device 10.

[0041] The interface 43 includes a communication device or a communication module configured to be able to communicate with the interface 13 of the audio processing device 10. The interface 43 may be configured identically to or similarly to the interface 13 of the audio processing device 10.

[0042] <Example of arrangement of components of audio processing system 1> In the audio processing system 1, the audio processing device 10, the sound collection device 20, the detection device 30, and the audio output device 40 may be configured separately. The audio processing device 10 may be configured integrally with at least one of the sound collection device 20, the detection device 30, or the audio output device 40. The sound collection device 20 may be configured integrally with at least one of the audio processing device 10, the detection device 30, or the audio output device 40. The detection device 30 may be configured integrally with at least one of the audio processing device 10, the sound collection device 20, or the audio output device 40. The audio output device 40 may be configured integrally with at least one of the audio processing device 10, the sound collection device 20, or the detection device 30.

[0043] The part where the detection device 30 is worn is not limited to the user's arm. The detection device 30 may be worn on the user's head, for example. The detection device 30 may be integrally configured with the audio output device 40 and worn on the user's ear. When the detection device 30 is worn on the user's head or ear, it can detect the movement or orientation of the user's head. The part where the detection device 30 is worn is also referred to as the wearing part. When the detection device 30 is worn on the user's arm, the wearing part is the user's arm. When the detection device 30 is worn on the user's head, the wearing part is the user's head.

[0044] The voice processing device 10 may be attached to the user's body, regardless of whether it is configured integrally with or separate from the detection device 30 or the voice output device 40. When the voice processing device 10 is not attached to the user's body, it may be carried by the user. The voice processing device 10 may be located in a location remote from the user. The voice processing device 10 may be configured in a cloud computing environment or an on-premise environment.

[0045] The sound collection device 20 may be attached to the user's body, regardless of whether it is configured integrally with the detection device 30 or the sound output device 40 or configured separately. If the sound collection device 20 is not attached to the user's body, it may be carried by the user. If the sound collection device 20 is not carried by the user, it may be located within a predetermined range from the user or at a location away from the user.

[0046] The detection device 30 may be configured to be wearable on the user's ear together with the sound output device 40. The sound collection device 20 may be configured to be wearable on the user's ear together with the sound output device 40. The sound collection device 20 may be configured to be wearable on the user's arm together with the detection device 30.

[0047] (Operation Example of Voice Processing System 1) Voices occurring around a user include voices occurring from multiple sound sources. In the voice processing system 1 according to the present disclosure, the voice processing device 10 outputs voices occurring from specific sound sources by emphasizing them so that the user can easily hear them. An operation example of the voice processing device 10 will be described below.

[0048] <Example of Operation for Assigning Weights According to the Direction from Which Sound is Generated> Sounds generated around the user include sounds generated from multiple directions. The sound processing device 10 emphasizes and outputs sounds generated from specific directions so that the user can easily hear them. Specifically, the sound processing device 10 separates the sounds generated around the user into multiple independent sounds. The sound processing device 10 identifies the direction indicated by the user based on the user's movement. The sound processing device 10 generates playback data in which weights are assigned to each of the separated independent sounds according to the direction from which each independent sound is generated, so as to emphasize the independent sounds generated in the direction indicated by the user.

[0049] The control unit 11 of the audio processing device 10 may execute an audio processing method including the steps of the flowchart illustrated in Fig. 3 in order to generate playback data in which the independent audio in the direction indicated by the user is emphasized. The audio processing method may be realized as an audio processing program executed by a processor constituting the control unit 11. The audio processing program may be stored in a non-transitory computer-readable medium.

[0050] The control unit 11 acquires audio data from the sound collection device 20 (step S1). Specifically, the sound collection device 20 detects audio generated around the user of the audio processing system 1 and outputs the audio data. It is assumed that the audio data includes audio generated from multiple sound sources located around the user. The audio generated from each sound source is independent from each other and is also referred to as independent audio. The independent audio may include, for example, human voices, background street sounds, or noise.

[0051] The control unit 11 separates the audio data into a plurality of independent voices (step S2). Separation of the audio data into a plurality of independent voices may be achieved, for example, by a neural network model or by various other techniques. The control unit 11 may include an independent voice extraction unit. The independent voice extraction unit may separate the audio data into a plurality of independent voices and extract data for each independent voice.

[0052] The control unit 11 identifies the direction from which each separated independent sound is generated and associates a direction with each separated independent sound (step S3). Identifying the direction from which each independent sound is generated is also called sound source localization. Sound source localization may be achieved, for example, by the MUSIC (Multiple Signal Classification) method or by various other techniques. The control unit 11 may include a sound source localization unit. The sound source localization unit may identify the direction from which each independent sound is generated. The control unit 11 may generate pairs of each independent sound and the direction from which each independent sound is generated. The control unit 11 may store information identifying pairs of each independent sound and a direction together with data of each independent sound in the storage unit 12. The control unit 11 may include an individual sound / direction linking unit. The individual sound / direction linking unit may generate pairs of each independent sound and a direction.

[0053] The independent voice extraction unit, the sound source localization unit, and each independent voice / direction linking unit are also collectively referred to as an acquisition candidate independent voice / direction pair calculation unit.

[0054] The control unit 11 acquires motion data from the detection device 30 (step S4). Specifically, the detection device 30 detects the user's movement as acceleration, angular velocity, or the like using the motion sensor 31, and outputs the detected movement as motion data. The motion data is data that represents the movement of the part of the body to which the motion sensor 31 is attached in terms of acceleration, angular velocity, or the like.

[0055] The control unit 11 identifies the selected direction based on the motion data (step S5). Specifically, when the detection device 30 is worn on the user's arm, the control unit 11 may identify, as the selected direction, the direction in which the user's arm is facing when the user's arm remains stationary for a predetermined period of time. The control unit 11 may set the predetermined period of time as appropriate. The control unit 11 may set the predetermined period of time based on a user input. The predetermined period of time may be set to, for example, 0.2 seconds, but is not limited to this and may be set to another value. The state in which the user's arm remains stationary for a predetermined period of time may correspond to a state in which the acceleration and angular velocity detected by the motion sensor 31 remain zero for a predetermined period of time, or may correspond to a state in which the acceleration and angular velocity remain below a stationary state determination threshold for a predetermined period of time. The stationary state determination threshold may be set as appropriate.

[0056] The control unit 11 may register a horizontal angle as a reference point in advance and identify the selected direction based on the difference in angle from the horizontal angle of the reference point. The horizontal angle corresponds to a 360-degree angle around the user in a horizontal plane. The orientation of the part of the body where the motion sensor 31 is attached is calculated based on angular velocity data acquired from when the part of the body was facing the horizontal angle as the reference point.

[0057] The reference point may be registered according to the orientation of the user, such as directly in front of the user. When setting the reference point according to the orientation of the user, the control unit 11 may register the reference point by having the user perform a predetermined movement and detecting the movement with the motion sensor 31. A method by which the control unit 11 registers the reference point by having the user perform a predetermined movement will be described later.

[0058] The reference point may be registered regardless of the user's orientation. The reference point may be registered as, for example, any of east-west, north-south, or various directions such as eight or sixteen directions. When registering the reference point as a direction, the control unit 11 may register the reference point based on the detection result of a direction detection device such as an electronic compass.

[0059] The control unit 11 sets a weight to be assigned to the independent voice for each direction (step S6). Specifically, the control unit 11 determines whether the direction associated with at least one independent voice matches the selected direction. That is, the control unit 11 determines whether an independent voice associated with the selected direction exists. The control unit 11 may determine that the direction associated with at least one independent voice matches the selected direction when the difference in angle between the direction associated with at least one independent voice and the selected direction is less than a direction determination threshold. That is, the control unit 11 may determine that an independent voice associated with the selected direction exists when the difference in angle between the direction associated with at least one independent voice and the selected direction is less than a direction determination threshold. The direction determination threshold may be set as appropriate.

[0060] When there is an independent voice associated with the selected direction, the control unit 11 sets a weight to be given to each independent voice so that the weight to be given to the independent voice associated with the selected direction is greater than the weight to be given to the independent voice associated with other directions.

[0061] The control unit 11 may set a weight to each independent voice so that the volume of the voice generated in the selected direction when played back is greater than the volume of the voice generated in other directions. The control unit 11 may set the weight to 0 to the independent voices in other directions so that only the independent voice associated with the selected direction can be selected and played back. In this way, even if the original volume of the voice generated in the selected direction is low, the user can easily hear the voice generated in the selected direction.

[0062] If there is no independent voice associated with the selected direction, the control unit 11 may generate playback data without weighting each independent voice.If there is no independent voice associated with the selected direction, the control unit 11 may increase the direction determination threshold and re-determine whether the direction associated with at least one independent voice matches the selected direction.

[0063] The control unit 11 generates playback data (step S7). Specifically, the control unit 11 generates playback data by combining the independent sounds from each direction with the weights set in step S6. After executing the procedure in step S7, the control unit 11 ends the execution of the procedure in the flowchart of FIG. 3.

[0064] The control unit 11 may output the playback data generated in the procedure of step S7 to the audio output device 40 and cause the audio output device 40 to play it. The audio output device 40 may output the playback data as audio to the user. By setting the weight assigned to the independent audio of directions other than the selected direction to a small value or 0, audio in which the independent audio of the selected direction is emphasized is output when the audio output device 40 generates the playback data.

[0065] The control unit 11 may reset the weight assigned to each independent voice in accordance with a change in the direction from which each independent voice is generated by repeating the execution of the procedure in the flowchart of Fig. 3. The control unit 11 may reset the weight assigned to each independent voice in accordance with a change in the selected direction. The control unit 11 may reset the weight assigned to each independent voice in accordance with a change in the direction from which each independent voice is generated and a change in the selected direction.

[0066] <Example of Operation for Assigning Weights According to Distance from Which Sound is Generated> Sound generated around the user may include sounds generated at multiple points located at different distances in the same direction. The sound processing device 10 may emphasize and output sounds generated at points located at specific distances in a specific direction so that the user can easily hear them. Specifically, the sound processing device 10 separates the sound generated around the user into multiple independent sounds. The sound processing device 10 identifies the direction selected by the user and the distance to the point selected by the user based on the user's movement. The sound processing device 10 assigns weights to each of the separated independent sounds according to the direction from which each independent sound is generated, so as to emphasize the independent sounds generated in the direction selected by the user and at the selected point, and generates playback data with weights further assigned according to distance. The point selected by the user is also referred to as the selected point.

[0067] The control unit 11 of the audio processing device 10 may execute an audio processing method including the steps of the flowchart illustrated in Fig. 3 in order to generate playback data in which the independent audio is emphasized. The audio processing method may be realized as an audio processing program executed by a processor constituting the control unit 11. The audio processing program may be stored in a non-transitory computer-readable medium.

[0068] The control unit 11 acquires audio data from the sound collection device 20 (step S11). The control unit 11 may acquire the audio data in a procedure that is the same as or similar to the procedure of step S1 in FIG.

[0069] The control unit 11 separates the audio data into a plurality of independent sounds (step S12). The control unit 11 may separate the audio data into a plurality of independent sounds using a procedure that is the same as or similar to the procedure of step S2 in FIG.

[0070] The control unit 11 identifies the direction and distance from which each separated independent sound is generated, and associates the direction and distance with each separated independent sound (step S13). The control unit 11 may perform the procedure for associating each independent sound with a direction in a manner that is the same as or similar to the procedure of step S3 in Fig. 3. The control unit 11 identifies the distance from the microphone 21 of the sound collection device 20 to the sound source of each independent sound, and further associates the distance with each independent sound.

[0071] The distance to the sound source can be calculated by various methods. For example, the distance to the sound source of each independent sound may be calculated based on the fact that the power of a sound detectable by the microphone 21 is inversely proportional to the distance between the sound source of that sound and the microphone 21. The control unit 11 may include an independent sound sound source distance calculation unit. Each independent sound sound source distance calculation unit may calculate the distance to the sound source of each independent sound. The control unit 11 may include an independent sound / distance linking unit. Each independent sound / distance linking unit may generate a pair of each independent sound and a distance.

[0072] The independent speech extraction unit, each independent speech sound source distance calculation unit, and each independent speech / distance linking unit are also collectively referred to as an acquisition candidate independent speech / distance pair calculation unit.

[0073] The control unit 11 acquires motion data from the detection device 30 (step S14). The control unit 11 may acquire the motion data in a procedure that is the same as or similar to the procedure of step S4 in FIG.

[0074] The control unit 11 identifies the selection direction based on the motion data (step S15). The control unit 11 may identify the selection direction using a procedure that is the same as or similar to the procedure of step S5 in FIG.

[0075] The control unit 11 sets the weights to be given to the independent sounds in each direction (step S16). The control unit 11 may set the weights using a procedure that is the same as or similar to the procedure in step S6 of FIG.

[0076] The control unit 11 sets the distance from the user to the selected point (step S17). The control unit 11 may set the distance from the user to the selected point based on the motion data. When the detection device 30 is worn on the user's arm, the control unit 11 may decrease or increase the distance to the selected point according to the movement of the user's arm. The control unit 11 may associate a specific gesture of the user's arm with an instruction to decrease the distance to the selected point or an instruction to increase the distance to the selected point. For example, the control unit 11 may associate a gesture of moving the user's arm away from the body toward the selected point, a gesture of flicking the back of the user's hand away from the body toward the selected point, or a gesture of flicking the user's palm away from the body toward the selected point with an instruction to increase the distance to the selected point. Furthermore, the control unit 11 may associate a gesture of moving the user's arm from the selected point toward the body, a gesture of flicking the back of the user's hand from the selected point toward the body, or a gesture of flicking the user's palm from the selected point toward the body with an instruction to decrease the distance to the selected point. The control unit 11 may detect the user's arm movement as a gesture using a pattern matching technique that determines whether motion data matches a template. The control unit 11 may use a machine learning model configured to input motion data and output an estimated gesture result. The control unit 11 may decrease or increase the distance to the selected point depending on the magnitude of the user's arm movement. For example, the control unit 11 may associate a user's arm movement greater than a predetermined threshold with an instruction to decrease the distance to the selected point. For example, the control unit 11 may associate a user's arm movement smaller than a predetermined threshold with an instruction to decrease the distance to the selected point.

[0077] The control unit 11 may set the distance to the selected point in accordance with various operations other than the motion data. The control unit 11 may set the distance to the selected point in accordance with an operation input from an input device. The input from the input device may be, for example, a button input, a touch sensor input, or a voice input.

[0078] The control unit 11 changes the weight of the independent voice for each distance in the selected direction (step S18). Specifically, the control unit 11 determines whether the distance associated with at least one independent voice matches the distance to the selected point. That is, the control unit 11 determines whether an independent voice exists that is associated with the distance to the selected point. The control unit 11 may determine that the distance associated with at least one independent voice matches the distance to the selected point when the difference between the distance associated with at least one independent voice and the distance to the selected point is less than the distance determination threshold. That is, the control unit 11 may determine that the independent voice exists that is associated with the distance to the selected point when the difference between the distance associated with at least one independent voice and the distance to the selected point is less than the distance determination threshold. The distance determination threshold may be set as appropriate.

[0079] When there is an independent voice associated with the distance to the selected point, the control unit 11 changes the weight given to each independent voice associated with the selected direction, that of the independent voices associated with distances other than the distance to the selected point, to a value smaller than the weight given to the independent voice associated with the distance to the selected point.

[0080] The control unit 11 may set a weight to be assigned to each independent voice so that the volume of the voice generated at the selected point when played back is louder than the volume of the voice generated at other points when played back. The control unit 11 may change the weight to be assigned to the independent voices associated with other distances to 0 so that only the independent voices associated with the distance to the selected point can be selected and played back. In this way, even if the original volume of the voice generated at the selected point is low, the user can easily hear the voice generated at the selected point.

[0081] When there is no independent voice associated with the distance to the selected point, the control unit 11 may generate playback data without changing the weight of each independent voice associated with the selected direction.When there is no independent voice associated with the distance to the selected point, the control unit 11 may increase the distance determination threshold and re-determine whether the distance associated with at least one independent voice matches the distance to the selected point.

[0082] The control unit 11 generates playback data (step S19). The control unit 11 may generate the playback data using a procedure that is the same as or similar to the procedure of step S7 in Fig. 3. After executing the procedure of step S19, the control unit 11 ends execution of the procedure of the flowchart in Fig. 4.

[0083] The control unit 11 may output the playback data generated in the procedure of step S19 to the audio output device 40 and cause the audio output device 40 to play it. The audio output device 40 may output the playback data as audio to the user. By setting the weighting of the independent audio of points other than the selected point to a small value or 0, audio in which the independent audio occurring at the selected point is emphasized when the audio output device 40 generates the playback data is output.

[0084] The control unit 11 may reset the weight assigned to each independent voice in accordance with a change in the direction or distance from which each independent voice is generated by repeating the execution of the procedure in the flowchart of Fig. 4. The control unit 11 may reset the weight assigned to each independent voice in accordance with a change in the selected direction or distance to the selected point. The control unit 11 may reset the weight assigned to each independent voice in accordance with a change in the direction or distance from which each independent voice is generated and a change in the selected direction or distance to the selected point.

[0085] <Registration of Reference Point Angle> As described above, the control unit 11 of the voice processing device 10 identifies a selection direction based on the user's movement and sets a weight to be assigned to an independent voice. The control unit 11 may execute a reference point angle registration method including the steps of the flowchart illustrated in FIG. 5 in order to register a reference point used to identify the selection direction. The reference point angle registration method may be included in the voice processing method. The reference point angle registration method may be realized as a reference point angle registration program executed by a processor constituting the control unit 11. The reference point angle registration program may be included in the voice processing program. The reference point angle registration program may be stored on a non-transitory computer-readable medium.

[0086] The control unit 11 starts to register the reference point angle (step S21). The reference point angle is a criterion used to identify the selection direction according to the direction of the user's arm.

[0087] The control unit 11 outputs an instruction to the user to stretch their arms out in front (step S22). The control unit 11 may output the instruction as sound from the speaker 41 of the audio output device 40. The audio output device 40 may include an audio instruction unit. The audio instruction unit of the audio output device 40 may output the instruction as sound. The control unit 11 may display the instruction on a display device of the interface 33 of the detection device 30. The control unit 11 may notify the user of the instruction by various means.

[0088] The control unit 11 determines whether the user's arm has been stationary for a predetermined time (step S23). Specifically, the control unit 11 acquires motion data that detects the movement of the user's arm from the detection device 30, and determines whether the movement of the user's arm has been stationary for a predetermined time. The predetermined time may be set to, for example, one second, or may be set to another value. The determination of whether the user's arm has been stationary for a predetermined time may be performed by an operation that is the same as or similar to the operation described as the operation for identifying the selection direction in step S4 of FIG. 3.

[0089] If the user's arm has not remained still for the specified time (step S23: NO), the control unit 11 returns to the procedure of step S22 and continues outputting instructions until the user's arm remains still for the specified time. After issuing an instruction once, the control unit 11 may repeatedly determine whether the user's arm has remained still for the specified time without issuing an instruction again.

[0090] If the user's arm remains stationary for a specified period of time (step S23: YES), the control unit 11 registers a first reference point (step S24). The first reference point is registered as the horizontal angle of the motion sensor 31 when the user's arm is extended in front. The horizontal angle is also referred to as the α angle. The control unit 11 may provisionally register the α angle of the motion sensor 31 when the user's arm is extended in front as 0 degrees. The control unit 11 may register the angle of the first reference point by storing it in the memory unit 12. The direction of the user's arm when registering the first reference point is also referred to as the first direction.

[0091] The detection device 30 may include a reference point angle information generation unit that generates information on an angle to be registered as the first reference point. The sound processing device 10 may include a reference point angle information registration unit that registers information on the angle of the first reference point.

[0092] The control unit 11 outputs an instruction to the user to stretch the arm to the right or left (step S25). The control unit 11 may output the instruction by the same or similar action as in the procedure of step S22.

[0093] The control unit 11 determines whether the user's arm has been stationary for a specified time (step S26). The control unit 11 may determine whether the user's arm has been stationary for a specified time by performing the same or similar action as the procedure of step S23.

[0094] If the user's arm has not remained still for the specified time (step S26: NO), the control unit 11 returns to the procedure of step S25 and continues outputting instructions until the user's arm remains still for the specified time. After issuing an instruction once, the control unit 11 may repeatedly determine whether the user's arm has remained still for the specified time without issuing an instruction again.

[0095] If the user's arm remains stationary for a predetermined period of time (step S26: YES), the control unit 11 registers a second reference point (step S27). The second reference point is expressed as a difference from the α angle at the first reference point. If the α angle of the motion sensor 31 when the user's arm is extended forward is set to 0 degrees in the first reference point setting procedure of step S24, the second reference point corresponds to the α angle itself. The control unit 11 calculates the α angle when the arm is extended to the right or left by integrating and square-norming the angular velocity detected by the motion sensor 31 while the arm transitions from a state in which the arm is extended forward to a state in which the arm is extended to the right or left, and registers the calculated α angle as the second reference point. The control unit 11 may register the angle of the second reference point by storing it in the memory unit 12. The direction of the user's arm when registering the second reference point is also referred to as the second direction.

[0096] When the detection device 30 includes a reference point angle information generation unit, the reference point angle information generation unit generates information on the angle to be registered as the second reference point. When the sound processing device 10 includes a reference point angle information registration unit, the reference point angle information registration unit registers information on the angle of the second reference point.

[0097] After executing the procedure of step S27, the control unit 11 ends the execution of the procedure of the flowchart in Fig. 5. The control unit 11 may set a virtual coordinate system with the first reference point as its origin, based on information about the angles of the first reference point and the second reference point.

[0098] <Determining Selection Direction> The control unit 11 of the voice processing device 10 may determine the selection direction based on the user's movement, using the α angle registered as a reference point as a reference. The control unit 11 may execute a selection direction determination method including the steps of the flowchart illustrated in FIG. 6 to determine the selection direction. The selection direction determination method may be included in a voice processing method. The selection direction determination method may be realized as a selection direction determination program executed by a processor constituting the control unit 11. The selection direction determination program may be included in the voice processing program. The selection direction determination program may be stored on a non-transitory computer-readable medium.

[0099] The control unit 11 acquires motion data from the detection device 30 (step S31). The audio processing device 10 may include a device angular velocity information acquisition unit. The device angular velocity information acquisition unit acquires motion data including angular velocity information from the detection device 30.

[0100] The control unit 11 calculates the direction of the user's arm based on the reference point angle information and the motion data (step S32). The direction of the user's arm is represented by an α angle. The control unit 11 may calculate the α angle representing the direction of the user's arm by integrating the angular velocity detected by the motion sensor 31 while the direction of the arm changes from the first reference point or the second reference point and taking the square norm. The control unit 11 may include a direction calculation unit that calculates the α angle.

[0101] The control unit 11 determines whether the user's arm has been stationary for the set time (step S33). The time for determining whether the user's arm has been stationary may be set to, for example, 0.2 seconds, but is not limited to this and may be set to another value. The state in which the user's arm has been stationary for the set time may correspond to a state in which the acceleration and angular velocity detected by the motion sensor 31 have remained zero for the set time, or may correspond to a state in which the acceleration and angular velocity have remained below the stationary determination threshold for the set time.

[0102] If the user's arm is not stationary for the set time (step S33: NO), the control unit 11 returns to the procedure of step S32 and repeats the procedure of calculating the direction of the user's arm and the procedure of determining whether the user's arm is stationary.

[0103] If the user's arm remains stationary for the set time (step S33: YES), the control unit 11 determines the direction of the stationary arm as the selected direction (step S34). The control unit 11 may include a direction determination unit that determines the selected direction. The direction calculation unit and direction determination unit are also collectively referred to as a selected direction calculation unit. After executing the procedure of step S34, the control unit 11 ends execution of the procedure of the flowchart in FIG. 6.

[0104] The control unit 11 may execute a selected point determination method including steps that are the same as or similar to the above-described selected direction determination method. The selected point determination method may include a step of detecting the user's arm movement as a gesture to set a distance in the selected direction, and a step of determining a point located at the set distance as the selected point. The control unit 11 may include a motion detection unit that detects the user's arm movement as a gesture. The control unit 11 may also include a distance calculation unit that calculates a distance based on the user's arm movement. The motion detection unit and the distance calculation unit are also collectively referred to as a selected distance calculation unit.

[0105] (Summary) As described above, the voice processing system 1 according to the present disclosure can separate voices generated around the user into independent voices, determine a selected direction based on the movement of the part of the user wearing the motion sensor 31, and emphasize and output the independent voices generated in the selected direction. The user can intuitively select a direction by moving the part of the user wearing the motion sensor 31. As a result, the user can easily select the voices they want to listen to.

[0106] Other Embodiments Other embodiments will be described below.

[0107] <Example of Functional Allocation> The sound processing device 10, the sound collection device 20, and the sound output device 40 may be configured as an integrated hearable device and worn by a user. In the hearable device, at least a part of the function of the control unit 11 of the sound processing device 10 may be realized by the control unit 22 of the sound collection device 20 or the control unit 42 of the sound output device 40.

[0108] Among the functions of the voice processing device 10, the function of determining the selection direction or the selection distance may be realized by the control unit 32 of the detection device 30. In this case, the detection device 30 may include a selection direction information transmission unit or a selection distance information transmission unit. The selection direction information transmission unit transmits the determined selection direction to the hearable device. The selection distance information transmission unit transmits the determined selection distance to the hearable device.

[0109] The hearable device may include a selected direction information receiving unit or a selected distance information receiving unit. The selected direction information receiving unit receives the selected direction determined by the detection device 30. The selected distance information receiving unit receives the selected distance determined by the detection device 30.

[0110] The hearable device may include a selected direction audio acquisition unit or a selected distance audio acquisition unit. The selected direction audio acquisition unit acquires independent audio paired with the selected direction based on the selected direction received by the selected direction receiving unit and the pair of independent audio and direction generated by the acquisition candidate independent audio / direction pair calculation unit. The selected distance audio acquisition unit acquires independent audio paired with the selected distance based on the selected distance received by the selected distance receiving unit and the pair of independent audio and distance generated by the acquisition candidate independent audio / distance pair calculation unit.

[0111] The hearable device may include a voice information transmitting unit that transmits information of an independent voice acquired by the selected direction voice acquiring unit or the selected distance voice acquiring unit as voice information to be heard by a user. The voice information corresponds to playback data.

[0112] <Time Required for Audio Processing> As described above, the audio processing device 10 processes the audio detected by the sound collection device 20 to generate playback data that emphasizes the audio in the direction selected by the user. The audio processing device 10 may process the audio occurring around the user in real time to generate the playback data.

[0113] The time required for the voice processing device 10 to process voice is a finite length greater than zero. The voice processing device 10 may shorten the voice processing time so that voice can be reproduced with a delay time that does not cause discomfort to the user. The allowable delay time may be set as appropriate. Even if the voice processing device 10 cannot process voice in real time when real time is interpreted as simultaneous, it can process voice in pseudo-real time by processing voice with an allowable delay time.

[0114] The sound processing device 10 may record the sound detected by the sound collection device 20, and generate playback data in which the sound from the selected direction is emphasized when playing back the recorded sound data.

[0115] <Example in which the motion sensor 31 is worn on the user's head> In the above-described embodiment, the detection device 30 is worn on the user's arm, and the selection direction is determined according to the orientation of the user's arm. When the detection device 30 is worn on the user's ear or head, the control unit 11 of the voice processing device 10 may determine the selection direction according to an α angle of the orientation of the user's head. The control unit 11 may register a first reference point based on motion data when the user's head is facing straight ahead, and may register a second reference point based on motion data when the user's head is facing right or left. The control unit 11 may calculate an α angle according to the orientation of the user's head to determine the selection direction.

[0116] The range in which the user's head can be rotated horizontally may be limited to a narrow angle. To enable a wide range of directions to be selected, the control unit 11 may determine the selected direction according to an angle obtained by multiplying the alpha angle of the user's head by a coefficient greater than 1. Even when determining the selected direction according to the direction of the user's arm, the control unit 11 may determine the selected direction according to the angle obtained by multiplying the alpha angle of the arm direction by a coefficient.

[0117] The control unit 11 may detect a pattern of movement of the user's head as a gesture, and set the distance in the selected direction according to the detected gesture.

[0118] <Relationship between Selected Direction and Direction from which Independent Sound is Generated> As described above, the control unit 11 of the sound processing device 10 identifies the direction from which an independent sound generated around the user is generated. The control unit 11 may identify the direction from which the independent sound is generated so that the selected direction determined as the direction from which the user's arm is facing matches the direction from which the independent sound that the user wants to listen to is generated. For example, the control unit 11 may recognize the direction of the directivity of the microphone 21 when the sound collection device 20 is worn on the user's ear based on the front of the user.

[0119] The detection device 30 and the sound collection device 20 may be worn on parts that move in different directions. For example, when the detection device 30 is worn on the user's arm and the sound collection device 20 is worn on the user's ear, the detection device 30 and the sound collection device 20 may move in different directions. If the detection device 30 registers the first reference point in the direction directly in front of the user and the user's head moves to the right or left, the relationship between the direction of the directivity of the microphone 21 that the control unit 11 previously recognized and the first reference point of the detection device 30 changes. In this case, the direction from which the independent sound that the user wants to hear is generated may not match the selected direction.

[0120] Therefore, the sound processing system 1 may further include a detection device 30 for detecting the orientation of the part of the body where the sound collection device 20 is worn, in addition to the detection device 30 for determining the user's selected direction. For example, when the sound collection device 20 is worn on the user's head, the control unit 11 can detect the orientation of the user's head, and thereby match the direction in which the independent sound that the user wants to hear is generated with the selected direction even if the user's head is facing in a direction other than forward.

[0121] When the detection device 30 and the sound collection device 20 are integrally configured, or when the detection device 30 and the sound collection device 20 are worn on a part where they move in the same direction, the detection device 30 and the sound collection device 20 move in the same direction regardless of the direction in which the part where they are worn faces. Therefore, the relationship between the direction of the directivity of the microphone 21 that the control unit 11 has recognized in advance and the first reference point of the detection device 30 does not change. As a result, regardless of the direction in which the part where they are worn faces, the direction in which the independent sound that the user wants to hear is generated can be matched with the selected direction.

[0122] When searching for an independent audio signal that the user wants to listen to, the user may not be able to clearly recognize, for example by visual inspection, the direction from which the independent audio signal that the user wants to listen to is coming from. As described above, the audio processing system 1 according to the present disclosure enables the user to easily select the audio signal that the user wants to listen to. Therefore, even if the direction from which the independent audio signal that the user wants to listen to is coming from does not match the selected direction, the function of easily selecting the audio signal that the user wants to listen to is not impaired.

[0123] Those skilled in the art may make various modifications and alterations to the contents of the present disclosure based on the present disclosure. Therefore, these modifications and alterations are within the scope of the present disclosure. For example, in each embodiment, each functional unit, each means, or each step may be added to other embodiments without logical inconsistency, or may be replaced with each functional unit, each means, or each step of other embodiments. Furthermore, in each embodiment, multiple functional units, each means, or each step may be combined into one or divided into separate units. Furthermore, the above-described embodiments of the present disclosure are not limited to faithful implementation of each of the described embodiments, and may be implemented by combining features or omitting some features as appropriate.

[0124] In one embodiment, (1) the audio processing device includes an acquisition unit that acquires audio data that captures audio occurring around a user and motion data that detects the user's movements, and a control unit that determines a selection direction to be selected by the user based on the motion data, and generates playback data by giving a greater weight to audio occurring in the selection direction among the audio data than to audio occurring in directions other than the selection direction.

[0125] (2) In the audio processing device described in (1) above, the control unit may set a weight to be given to the audio generated in each direction so that the volume of the audio generated in the selected direction when played back is louder than the volume of the audio generated in directions other than the selected direction when played back.

[0126] (3) In the audio processing device described in (2) above, the control unit may generate, as the playback data, data in which only audio generated in the selected direction is selected from the audio data.

[0127] (4) In the audio processing device described in any one of (1) to (3) above, the control unit may set a distance to a selected point located in the selected direction based on the motion data, and may assign a greater weight to audio occurring at the selected point among audio occurring in the selected direction than to audio occurring at points other than the selected point.

[0128] (5) In the audio processing device described in (4) above, the control unit may set a weight to be given to the audio generated at each point in the selected direction so that the volume of the audio generated at the selected point when played back is louder than the volume of the audio generated at points other than the selected point when played back.

[0129] (6) In the audio processing device described in (5) above, the control unit may generate, as the playback data, data in which only audio occurring at the selected point is selected from the audio data.

[0130] (7) In the audio processing device according to any one of (1) to (6), the acquisition unit may acquire, as the motion data, a first reference point corresponding to an α angle when a wearing part on which a detection device that generates the motion data is worn faces a first direction, and a second reference point corresponding to an α angle when the wearing part faces a second direction. The control unit may determine the selected direction based on the first reference point, the second reference point, and the motion data.

[0131] In one embodiment, (8) the audio processing system includes a sound collection device that collects sounds occurring around a user and generates audio data; a detection device that detects the user's movements and outputs the motion data; an audio processing device that acquires the audio data from the sound collection device, acquires the motion data from the detection device, determines a selection direction to be selected by the user based on the motion data, and generates playback data by giving a greater weight to sounds occurring in the selection direction among the audio data than to sounds occurring in directions other than the selection direction; and an audio output device that outputs the playback data to the user as audio.

[0132] (9) In the audio processing system described in (8) above, the detection device may be configured to be wearable on the ear of the user together with the audio output device.

[0133] (10) In the audio processing system described in (8) or (9) above, the sound collection device may be configured to be wearable on the ear of the user together with the audio output device.

[0134] (11) In the sound processing system described in (8) above, the sound collection device may be configured to be wearable on the arm of the user together with the detection device.

[0135] In one embodiment, (12) an audio processing method includes an audio processing device acquiring audio data that captures audio occurring around a user and motion data that detects the movement of the user, determining a selection direction to be selected by the user based on the motion data, and generating playback data by giving a greater weight to audio occurring in the selection direction in the audio data than to audio occurring in directions other than the selection direction.

[0136] In one embodiment, (13) the audio processing program causes a processor to acquire audio data that captures audio occurring around the user and motion data that detects the user's movements, determine a selection direction to be selected by the user based on the motion data, and generate playback data by giving a greater weight to audio occurring in the selection direction among the audio data than to audio occurring in directions other than the selection direction.

[0137] 1 Audio processing system 10 Audio processing device (11: control unit, 12: memory unit, 13: interface) 20 Sound collection device (21: microphone, 22: control unit, 23: interface) 30 Detection device (31: motion sensor, 32: control unit, 33: interface) 40 Audio output device (41: speaker, 42: control unit, 43: interface)

Claims

1. An acquisition unit that acquires voice data obtained by collecting voices generated around a user and motion data obtained by detecting the user's motion, and a control unit that determines a selection direction selected by the user based on the motion data and generates reproduction data by setting a weight given to the voice generated in the selection direction in the voice data to be greater than a weight given to the voice generated in other directions than the selection direction. A voice processing apparatus comprising:

2. The voice processing apparatus according to claim 1, wherein the control unit sets a weight given to the voice generated in each direction so that a volume when the voice generated in the selection direction is reproduced is greater than a volume when the voice generated in other directions than the selection direction is reproduced.

3. The voice processing apparatus according to claim 2, wherein the control unit generates, as the reproduction data, data obtained by selecting only the voice generated in the selection direction from the voice data.

4. The voice processing apparatus according to any one of claims 1 to 3, wherein the control unit sets a distance to a selection point located in the selection direction based on the motion data, and sets a weight given to the voice generated at the selection point in the voice generated in the selection direction to be greater than a weight given to the voice generated at other points than the selection point.

5. The voice processing apparatus according to claim 4, wherein the control unit sets a weight given to the voice generated at each point in the selection direction so that a volume when the voice generated at the selection point is reproduced is greater than a volume when the voice generated at other points than the selection point is reproduced.

6. The voice processing apparatus according to claim 5, wherein the control unit generates, as the reproduction data, data obtained by selecting only the voice generated at the selection point from the voice data.

7. The voice processing apparatus according to any one of claims 1 to 6, wherein the acquisition unit acquires, as the motion data, a first reference point corresponding to an α angle when a wearing part on which a detection device that generates the motion data is worn faces a first direction, and a second reference point corresponding to the α angle when the wearing part faces a second direction, and the control unit determines the selection direction based on the first reference point, the second reference point, and the motion data.

8. A voice processing system comprising: a sound collection device that generates voice data by collecting voices generated around a user; a detection device that detects the movement of the user and outputs it as motion data; an audio processing device that acquires the voice data from the sound collection device, acquires the motion data from the detection device, determines a selection direction selected by the user based on the motion data, and generates reproduction data by making the weight given to the voice generated in the selection direction among the voice data larger than the weight given to the voice generated in other directions of the selection direction; and a voice output device that outputs the reproduction data as voice to the user.

9. The voice processing system according to claim 8, wherein the detection device is configured to be wearable on the user's ear together with the voice output device.

10. The voice processing system according to claim 8 or 9, wherein the sound collection device is configured to be wearable on the user's ear together with the voice output device.

11. The voice processing system according to claim 8, wherein the sound collection device is configured to be wearable on the user's arm together with the detection device.

12. A voice processing method including: acquiring voice data obtained by collecting voices generated around a user and motion data obtained by detecting the movement of the user; determining a selection direction selected by the user based on the motion data; and generating reproduction data by making the weight given to the voice generated in the selection direction among the voice data larger than the weight given to the voice generated in other directions of the selection direction.

13. A voice processing program for causing a processor to execute: acquiring voice data obtained by collecting voices generated around a user and motion data obtained by detecting the movement of the user; determining a selection direction selected by the user based on the motion data; and generating reproduction data by making the weight given to the voice generated in the selection direction among the voice data larger than the weight given to the voice generated in other directions of the selection direction.

Citation Information

Patent Citations

  • Mixing device, method, and program

    WO2006137400A1