Sound processing method and sound processing device
The sound processing method enhances the clarity of the primary speaker's voice in remote conferences by separating and emphasizing their direct sound component, addressing the challenge of distinguishing main conversation from background chatter.
Patent Information
- Application Number
- JP2024031212
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-01
- Publication Date
- 2025-09-11
AI Technical Summary
Existing sound processing devices struggle to distinguish between the voice of the primary speaker and secondary speakers during remote conferences, leading to difficulty in identifying the main conversation amidst background chatter.
A sound processing method that separates the voice of the primary speaker from secondary speakers using location information, emphasizing the direct sound component of the primary speaker's voice and adjusting sound levels based on speaker position.
Improves clarity of the primary speaker's speech by enhancing the direct sound component and reducing the prominence of secondary speakers' voices, making it easier to follow the main conversation.
Smart Images

Figure 2025133324000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a sound processing method and a sound processing device used in remote conferences and the like. [Background technology]
[0002] Patent Document 1 describes an information processing device that makes the speaking user's voice easier to hear by localizing the sound image of the speaking user's voice and sound images of background sounds, sound effects, etc. at predetermined positions. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] International Publication No. 2022 / 54900 Summary of the Invention [Problem to be solved by the invention]
[0004] However, the information processing device of Patent Document 1 improves the ease of listening to the voice of the speaker by changing the position where the sound image of the sound signal is localized. Therefore, the information processing device of Patent Document 1 does not perform sound processing by distinguishing between the voice of the speaker that is necessary for the conference (main conversation) and voice that is unnecessary for the conversation (utterances other than the main conversation). For this reason, users of Patent Document 1 may find it difficult to distinguish between the main conversation and utterances other than the main conversation (small talk, backchannels, etc.) during a remote conference.
[0005] An object of the present invention is to provide a sound processing method that improves the clarity of what the main speaker is saying in a remote conference or the like. [Means for solving the problem]
[0006] A sound processing method according to one embodiment of the present invention acquires a sound signal including the voice of a primary first speaker and a second speaker other than the first speaker, acquires location information of the first speaker, separates the voice of the first speaker and the voice of the second speaker based on the location information of the first speaker, and emphasizes the direct sound component of the voice of the first speaker. [Effects of the Invention]
[0007] According to the sound processing method of the present invention, it is possible to improve the ease with which the speech of the main speaker can be heard. [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a block diagram showing an example of the configuration of a sound processing system 1 according to a first embodiment. [Figure 2] 1 is an example of a three-dimensional schematic diagram of a room in which a sound processing system 1 according to a first embodiment is installed. [Figure 3] 1 is a block diagram showing an example of the configuration of a sound processing device 10 according to a first embodiment. [Figure 4] 1 is a functional block diagram illustrating an example of a sound processing method according to a first embodiment. [Figure 5] 4 is a flowchart illustrating an example of a sound processing method according to the first embodiment. [Figure 6] 3A and 3B are schematic diagrams showing the relationship between a direct sound signal and an indirect sound signal of an audio signal. [Figure 7] FIG. 10 is a functional block diagram showing an example of a sound processing method according to Modification 1. [Figure 8] FIG. 10 is a functional block diagram showing an example of a sound processing method according to Modification 2. [Figure 9] FIG. 11 is a functional block diagram showing an example of a sound processing method according to Modification 3. [Figure 10] FIG. 11 is a functional block diagram showing an example of a sound processing method according to Modification 4. [Figure 11] FIG. 10 is a diagram illustrating an example of a conversation history. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, a signal processing device according to an embodiment of the present invention will be described with reference to the drawings. In each drawing, the same parts are assigned the same reference numerals. From Modification 1 onwards, a description of matters common to the first embodiment will be omitted, and only the differences will be described. In particular, similar actions and effects resulting from similar configurations will not be mentioned in each embodiment.
[0010] First Embodiment FIG. 1 is a block diagram showing an example of the configuration of a sound processing system 1 according to a first embodiment of the present invention.
[0011] 1, the sound processing system 1 is made up of a sound processing device 10 and a personal computer (PC) 20. The sound processing device 10 is connected to the PC 20. The PC 20 is connected to a remote PC via a communication network 30. As a result, the PC 20 and the remote PC communicate data with each other via the network 30.
[0012] FIG. 2 is an example of a three-dimensional schematic diagram of a room in which the sound processing system 1 according to the first embodiment of the present invention is installed. As an example, the sound processing device 10 is installed on a wall surface of the room. Four microphones 15A-15D are installed at the four corners of a desk. A speaker 16 is installed on the sound processing device 10. Four participants A and D are located around the desk. A PC 20 is installed on the desk in front of participant A. From here on, the space in which the sound processing device 10 is installed will be referred to as the near-end side.
[0013] FIG. 3 is a block diagram showing an example of the configuration of the sound processing device 10 according to the first embodiment of the present invention.
[0014] The sound processing device 10 includes an audio interface (I / F) 101, a processor 102, a memory 103, and a communication interface (I / F) 104. The sound processing device 10 is also connected to four microphones 15A-15D and a speaker 16 via the audio I / F 101.
[0015] Memory 103 is a storage medium that stores an operation program for processor 102. Processor 102 reads the operation program from memory 103 and performs various operations. Note that the program does not have to be stored in memory 103. For example, the program may be stored in a storage medium of an external device such as a server. In this case, processor 102 simply reads the program from the server and executes it each time.
[0016] The processor 102 receives sound signals acquired by the four microphones 15A-15D via the audio I / F 101. The processor 102 performs predetermined processing on the sound signals acquired by the four microphones 15A-15D and outputs the sound signals to the communication I / F 104.
[0017] In the first embodiment, the number of microphones is four, but the number of microphones is not limited to four. The number of microphones may be one, or two or more.
[0018] The communication I / F 104 is a communication I / F such as USB, HDMI (registered trademark), or Bluetooth (registered trademark). The communication I / F 104 is connected to the PC 20 via, for example, USB. The communication I / F 104 outputs the sound signal acquired from the processor 102 to the PC 20. The PC 20 transmits the sound signal via the network 30 to a PC installed in a remote location.
[0019] The PC 20 receives a sound signal from a PC installed in a remote location. The PC 20 transmits the sound signal to the sound processing device 10 via the communication I / F 104. The processor 102 outputs the sound signal received via the communication I / F 104 to the speaker 16. The speaker 16 emits a sound based on the sound signal received from the processor 102.
[0020] With the above-described configuration, a user of the sound processing device 10 can hold a remote conference with a user in a remote location.
[0021] (Specific details of sound signal processing) The sound processing device 10 processes sound signals in a remote conference as follows: From here on, a sound processing method will be described in which sound signals acquired in a near-end space where the sound processing device 10 is installed are output to a far-end space.
[0022] Fig. 4 is a functional block diagram showing an example of a sound processing method according to the first embodiment, and Fig. 5 is a flowchart showing an example of a sound processing method according to the first embodiment.
[0023] 4, the processor 102 of the sound processing device 10 is composed of a sound signal processing unit 200 and an adder 210. The sound signal processing unit 200 is functionally composed of a sound signal acquisition unit 201, a position detection unit 202, a separation unit 203, and a first sound signal processing unit 204. The sound signal acquisition unit 201, the position detection unit 202, the separation unit 203, and the first sound signal processing unit 204 are programs executed by the processor 102. The first sound signal processing unit 204 is an example of a voice enhancement processing unit of the present invention.
[0024] The sound signal acquisition unit 201 acquires sound signals from the four microphones 15A-15D (S001).
[0025] Next, the position detection unit 202 acquires the position information of the first speaker (S002). Note that the first speaker here refers to a participant making a statement that constitutes the main conversation in the remote conference. The position detection unit 202 detects the voice of the first speaker based on, for example, sound signals acquired by the microphones 15A-15D. Specifically, the position detection unit 202 compares the levels of the sound signals acquired by the microphones 15A-15D and detects the sound signal with the highest level as the voice of the first speaker. The position detection unit 202 acquires the position information of the first speaker by calculating the arrival time difference of the sound signals of the voice of the first speaker acquired by the microphones 15A-15D, respectively.
[0026] Next, the separation unit 203 separates the voice of the first speaker from the voice of the second speaker based on the position information of the first speaker (S003). Note that the second speaker here refers to a participant in the remote conference who is making remarks unrelated to the main conversation (such as casual conversation or nodding along). In other words, the second speaker is a speaker other than the first speaker. For example, the separation unit 203 separates a sound signal generated from a direction that matches the position information of the first speaker from the sound signals acquired by the microphones 15A-15D. Specifically, the separation unit 203 separates the sound signals acquired by the microphones 15A-15D into the voice of the first speaker and the voice of the second speaker by, for example, linear filtering.
[0027] Next, the separation unit 203 separates the direct sound component from the voice of the first speaker. FIG. 6 is a schematic diagram showing the relationship between the direct sound signal and the indirect sound signal of the voice signal. The horizontal axis of the graph shown in FIG. 6 represents time, and the vertical axis represents amplitude. For example, when the separation unit 203 detects a level where the absolute value of the amplitude exceeds a predetermined threshold (the peak of the direct sound signal from the first speaker to microphones 15A-15D), the separation unit 203 defines the sound signal from the peak of the direct sound signal ±5 msec as the direct sound component. The separation unit 203 defines the part of the sound signal of the voice of the first speaker acquired by microphones 15A-15D excluding the direct sound component as the indirect sound component. Of course, the way in which the direct sound component and the indirect sound component are defined is not limited to this example. For example, the separation unit 203 may define the component from 0 to 50 msec as the direct sound component, and the component from 50 msec onwards as the indirect sound component. Alternatively, the separation unit 203 may define the components up to the reference time point at which the sound signal of the first speaker's voice reaches a predetermined level (for example, 1 / 2) of the peak level of the direct sound signal as the direct sound component, and the components after the reference time point as the indirect sound component. Alternatively, the separation unit 203 may include an FIR filter for canceling the acoustic characteristics of the space on the near-end side, and separate only the direct sound component from the sound signal of the first speaker's voice. In this case, the separation unit 203 acquires the acoustic characteristics of the space on the near-end side in advance as an impulse response, and calculates an inverse filter of the impulse response. The separation unit 203 separates the direct sound component by convolving the inverse filter with the input sound signal of the first speaker's voice.
[0028] Next, the first sound signal processing unit 204 emphasizes the direct sound component of the first speaker's voice (S004). The first sound signal processing unit 204 emphasizes the direct sound component by, for example, performing frequency characteristic adjustment processing (equalizer processing) that changes the frequency characteristics of the direct sound component of the first speaker's voice. The farther the speaker's voice is, the more attenuated it becomes. Furthermore, the farther the speaker's voice is, the more attenuated the high-frequency components of the speaker's voice are compared to the low-frequency components of the speaker's voice. Therefore, based on the position information of the first speaker, the greater the value of the distance to the microphones 15A-15D, the higher the level of the high-frequency band of the direct sound component of the first sound signal may be set.
[0029] Next, adder 210 adds the sound signal of the first speaker's voice processed by first sound signal processing unit 204 and the sound signal of the second speaker's voice separated by separation unit 203, and outputs the result to the far-end side via network 30.
[0030] As described above, the sound processing system 1 according to the first embodiment can adjust the output of the sound signal of the first speaker's voice to an optimal sound quality in accordance with the progress of the participants' conversation during a remote conference. More specifically, even when multiple participants are speaking simultaneously, the sound processing system 1 can detect the first speaker who is making a statement that constitutes the main conversation based on the sound signals acquired by the microphones 15A-15D, and can emphasize only the direct sound component of the first speaker's voice and output it to the far-end side. In other words, even when participants other than the first speaker are making statements unrelated to the main conversation (such as casual conversation or backchanneling), the sound signals of the voices of participants other than the first speaker are not emphasized. As described above, the user of the sound processing system 1 according to the first embodiment can easily identify the main conversation during a remote conference and can enjoy a new customer experience of being able to hear the statements of the speakers that constitute the main conversation with high sound quality.
[0031] Note that the processing of the direct sound component of the first speaker's voice in the present invention is not limited to the example described above. For example, the first sound signal processing unit 204 may emphasize the direct sound component of the first speaker's voice by performing a gain adjustment process on the direct sound component of the first speaker's voice. Specifically, the first sound signal processing unit 204 may set a gain such that the output power of the direct sound component of the first speaker's voice is equal to or greater than a predetermined value. Furthermore, the first sound signal processing unit 204 may set a gain based on the position information of the first speaker such that the level of the direct sound component of the first speaker's voice increases as the distance to the microphones 15A-15D increases.
[0032] Furthermore, the microphones 15A-15D in the present invention may be directional microphones. For example, the microphones 15A-15D each point their directivity in the direction of the four participants A, A, and B. Specifically, the microphone 15A points its directivity toward the position of participant A. The microphone 15B points its directivity toward the position of participant B. The microphone 15C points its directivity toward the position of participant C. The microphone 15D points its directivity toward the position of participant D. In this case, the position detection unit 202 may acquire position information of the first speaker based on the pickup ratio of sound signals acquired by the directional microphones 15A-15D. Specifically, the position detection unit 202 detects the acquisition time of each sound signal whose level exceeds a predetermined threshold based on the sound signals acquired by the microphones 15A-15D. The position detection unit 202 acquires the position of the microphone that picked up the sound signal with the longest acquisition time as the position information of the first speaker.
[0033] Furthermore, the first sound signal processing unit 204 may further process the indirect sound component of the first speaker's voice in addition to the direct sound component of the first speaker's voice. For example, the first sound signal processing unit 204 performs a process (reduction process) to reduce the indirect sound component of the first speaker's voice. Specifically, the first sound signal processing unit 204 reduces the gain of the indirect sound signal of the first speaker's voice or sets it to zero.
[0034] As a result, in the sound processing system 1 of the first embodiment, the ratio of indirect sound components to direct sound components of the voice of the first speaker is reduced, and the clarity of the voice of the first speaker can be further improved.
[0035] Variation 1 FIG. 7 is a functional block diagram showing an example of a sound processing method according to Modification 1. The sound processing device 10a of Modification 1 differs from the sound processing device 10 in that it further has a processing function for a direct sound signal and an indirect sound signal of a second speaker, in addition to the direct sound signal and the indirect sound signal of the first speaker's voice. More specifically, the sound processing device 10a functionally further includes a second sound signal processing unit 205. The second sound signal processing unit 205 is a program executed by the processor 102. The second sound signal processing unit 205 is an example of a voice enhancement processing unit of the present invention. The method for separating the sound signal of the second speaker's voice into a direct sound component and an indirect sound component is similar to the method for separating the sound signal of the first speaker's voice into a direct sound component and an indirect sound component, and therefore a description thereof will be omitted.
[0036] The second sound signal processing unit 205 adjusts the levels of the direct sound component and the indirect sound component of the second speaker's voice, for example. Specifically, the second sound signal processing unit 205 performs a reduction process on the direct sound component of the second speaker's voice, thereby adjusting the indirect sound component of the second speaker's voice so that it becomes dominant over the direct sound component of the second speaker's voice. For example, the second sound signal processing unit 205 reduces the gain of the direct sound component of the second speaker's voice. Alternatively, the second sound signal processing unit 205 may perform an adjustment process on the indirect sound component of the second speaker's voice, thereby adjusting the indirect sound component of the second speaker's voice so that it becomes dominant over the direct sound component of the second speaker's voice. For example, the second sound signal processing unit 205 increases the gain of the indirect sound component of the second speaker's voice.
[0037] Furthermore, the second sound signal processing unit 205 may determine how to process the direct sound component and indirect sound component of the second speaker depending on the progress of the main conversation. Specifically, when all sound signals collected by the four microphones 15A-15D are below a predetermined threshold, the position detection unit 202 determines that this is a section where the main conversation is not progressing (hereinafter referred to as a chat section), and determines that the collected sound signal is voice from the second speaker chatting. During the chat section, the second sound signal processing unit 205 makes an adjustment, for example, to extend the decay time of the reverberation sound of the indirect sound component of the second sound signal.
[0038] Generally, the longer the decay time of reverberation, the lower the clarity of the speech. As a result, the far-end participant perceives the speech (chat, etc.) of the second speaker on the near-end as being far away, and hears the speech as unclear.
[0039] In this way, the sound processing method of Variation 1 can adjust the sound output to an optimal level depending on the speech situations of the participants. More specifically, while the first speaker is speaking, the sound is adjusted so that the direct sound component of the first speaker's voice is more dominant than the direct sound component of the second speaker. Therefore, the ratio of the first speaker's voice to the second speaker's voice increases, making the first speaker's voice clearer and making it easier to distinguish the main conversation. Furthermore, during chat periods, the indirect sound component of the second speaker's voice is emphasized. Therefore, by increasing the ratio of the indirect sound component to the direct sound component of the second speaker's voice, the clarity of the second speaker's speech can be reduced. As a result, during chat periods, participants hear the second speaker's chatter or other sounds as unclear speech occurring from a distance, making it easy to distinguish that the speech is unrelated to the main conversation.
[0040] Variation 2 8 is a functional block diagram showing an example of a sound processing method according to Modification 2. The sound processing device 10b of Modification 2 differs from the sound processing device 10a in that a camera 17 is further connected and functionally includes a video signal acquisition unit 206. The video signal acquisition unit 206 is a program executed by the processor 102.
[0041] The sound processing device 10b acquires the position information of the first speaker based on the video signal received from the camera 17 in the process of step S002 in FIG.
[0042] The camera 17 captures a facial image of the participant AD on the near end side, and transmits a video signal to the video signal acquisition unit 206 .
[0043] The video signal acquiring unit 206 analyzes the video signal to detect the first speaker. More specifically, the video signal acquiring unit 206 detects the mouths of the participants from the video signal, for example, and analyzes the movement of the mouths to detect the first speaker. The video signal acquiring unit 206 transmits the position information of the detected first speaker to the position detecting unit 202.
[0044] In this way, in the sound processing method of the second modification, the video signal received from the camera 17 is analyzed to obtain the position information of the first speaker.
[0045] Note that position detection unit 202 may obtain the position information of the first speaker based on both the video signal and the sound signals picked up by microphones 15A-15D.
[0046] This allows the sound processing device 10b to acquire the position information of the first speaker with higher accuracy.
[0047] Variation 3 9 is a functional block diagram showing an example of a sound processing method according to Modification 3. The sound processing device 10c according to Modification 3 further includes a sound image localization processing unit 207. The sound processing device 10c of Modification 3 differs from the sound processing device 10b in that it has a processing function for localizing the sound images of the direct sound signal and indirect sound signal of the voice of the first speaker based on the position information of the first speaker. The sound image localization processing unit 207 is a program executed by the processor 102.
[0048] Camera 17 performs framing processing, for example, by panning, tilting, or zooming, so that only the first speaker appears in the image. Alternatively, camera 17 performs framing processing so that all participants on the near-end side appear in the image in addition to the first speaker. The image signal captured by camera 17 is preferably displayed on a display unit (not shown) on the far-end side.
[0049] The sound image localization processor 207 performs sound image localization processing on the direct sound component of the voice of the first speaker based on the position information of the first speaker. The sound image localization processing is a process of localizing a sound image so that the sound output from a speaker (not shown) installed in the space on the far end side appears as if it were generated at a predetermined position. The sound image localization processor 207 realizes the sound image localization processing by convolving a head-related transfer function with the direct sound component of the voice of the first speaker. The sound image localization processor 207 acquires the head-related transfer function from, for example, a memory provided in the sound processing device 10c, a network, an external storage medium, or the like.
[0050] Here, the head-related transfer function will be explained in more detail. The head-related transfer function is a function that expresses the transfer characteristics from the position of a sound source to the user's head (specifically, the user's left ear and right ear). There are two head-related transfer functions: one from the sound source to the left ear and one to the right ear. The sound image localization processing unit 207 convolves the sound signals corresponding to the L channel and R channel (direct sound components of the first speaker's voice) with the head-related transfer function to the left ear and the head-related transfer function to the right ear, respectively.
[0051] Furthermore, sound image localization processing unit 207 may perform panning processing on the direct sound component of the voice of the first speaker based on the position information of the first speaker. Panning processing is processing for adjusting the level ratio of the sound signal supplied to a speaker (not shown) installed on the far end side, and can change the sound image localization position of the sound signal in the horizontal direction.
[0052] In this way, the sound processing method of Variation 3 makes it possible to change the position where the sound image of the direct sound component of the voice of the first speaker output to the far-end side is localized. More specifically, by convolving the direct sound component of the voice of the first speaker with a head-related transfer function, the participant on the far-end side can perceive the voice of the first speaker on the near-end side as if it were coming from the actual location of the first speaker.
[0053] Furthermore, when a video signal is displayed on a display on the far-end side, the far-end participants perceive the sound component as coming directly from the position of the first speaker displayed on the display. Furthermore, when a video signal that has been framed for only the first speaker is displayed on a display on the far-end side, the far-end participants can easily identify the main speaker.
[0054] Variation 4 Fig. 10 is a functional block diagram showing an example of a sound processing method according to Modification 4. A sound processing device 10d according to Modification 4 further includes a conversation history recording unit 208. The sound processing device 10d of Modification 4 differs from the sound processing device 10c in that it has a function of analyzing the conversation history. The conversation history recording unit 208 is a program executed by the processor 102. Fig. 11 is a diagram showing an example of a conversation history.
[0055] The conversation history recording unit 208 records the conversation history in chronological order. The conversation history here refers to information indicating who the first speaker was. For example, in the example of FIG. 11, near-end participant A speaks from time t1 to time t2, and far-end participant E (not shown) speaks from time t2 to time t3. Near-end participant B speaks from time t3 to time t4, and far-end participant F (not shown) speaks from time t4 to time t5. Near-end participant C speaks from time t5 to time t6, and far-end participant G (not shown) speaks from time t6 to time t7. Near-end participant D speaks from time t7 to time t8, and far-end participant H (not shown) speaks from time t8 to time t9. Note that recording time information is not essential. The conversation history recording unit 208 may simply record an identifier indicating who the first speaker was and the speaker's order (number).
[0056] The conversation history recording unit 208 calculates the total speaking time as the first speaker for each participant based on the conversation history, and stores it as an index of the degree of participation in the conversation.
[0057] The position detection unit 202 acquires information indicating the degree of participation in the conversation for each participant from the conversation history recording unit 208. The position detection unit 202 acquires the position information of the first speaker based on both the position information of the first speaker acquired from the video signal acquisition unit 206 and the information acquired from the conversation history recording unit 208.
[0058] In this way, in the sound processing method according to the fourth modification, the accuracy of acquiring the location information of the first speaker can be further improved by analyzing the participation level in the conference for each participant based on the conversation history.
[0059] Finally, the description of the present embodiment should be considered to be illustrative in all respects and not restrictive. The scope of the present invention is defined not by the above-described embodiments but by the claims. Furthermore, the scope of the present invention is intended to include all modifications within the meaning and scope of the claims. [Explanation of symbols]
[0060] 1: sound processing system, 10, 10a, 10b, 10c, 10d: sound processing device, 20: PC, 30: network, 15A to 15D: microphone, 16: speaker, 17: camera, 101: audio I / F, 102: processor, 103: memory, 104: communication I / F, 200: sound signal processing unit, 210: adder, 201: sound signal acquisition unit, 202: position detection unit, 203: separation unit, 204: first sound signal processing unit, 205: second sound signal processing unit, 206: video signal acquisition unit, 207: sound image localization processing unit, 208: conversation history recording unit
Claims
1. Acquiring a sound signal including a voice of a primary first speaker and a voice of a second speaker other than the primary first speaker; acquiring location information of the first speaker; Separating the voice of the first speaker and the voice of the second speaker based on the position information of the first speaker; Emphasizing the direct sound component of the voice of the first speaker; Sound processing methods.
2. Separating the voice of the first speaker into the direct sound component and the indirect sound component; performing a process of reducing the indirect sound component; The sound processing method according to claim 1 .
3. performing a frequency characteristic adjustment process for the direct sound component; The sound processing method according to claim 2 .
4. Separating the speech of the second speaker into a direct sound component and an indirect sound component; performing a process of reducing the direct sound component of the second speaker; The sound processing method according to claim 3 .
5. performing an adjustment process for the indirect sound component of the second speaker; The sound processing method according to claim 4.
6. determining location information of the first speaker based on the sound signal; The sound processing method according to any one of claims 1 to 5.
7. acquiring a video signal including facial images of the first speaker and the second speaker; determining position information of the first speaker based on the video signal; The sound processing method according to any one of claims 1 to 5.
8. performing framing processing for the first speaker based on position information of the first speaker; The sound processing method according to claim 7.
9. performing a localization process for localizing a direct sound component of the first speaker based on the position information of the first speaker; The sound processing method according to any one of claims 1 to 5.
10. a sound signal acquisition unit that acquires a sound signal including a voice of a first speaker and a second speaker other than the first speaker; a position detection unit for acquiring position information of the first speaker; a separation unit that separates the voice of the first speaker and the voice of the second speaker based on position information of the first speaker; a voice enhancement processing unit that enhances a direct sound component of the voice of the first speaker. Sound processing device.
11. the separation unit separates the voice of the first speaker into the direct sound component and the indirect sound component; the speech emphasis processing unit performs a process of reducing the indirect sound component. The sound processing device according to claim 10.
12. the voice emphasis processing unit performs a frequency characteristic adjustment process on the direct sound component. The sound processing device according to claim 11 .
13. the separation unit separates the voice of the second speaker into a direct sound component and an indirect sound component; the speech enhancement processing unit performs a process of reducing the direct sound component of the second speaker. The sound processing device according to claim 12.
14. the speech enhancement processing unit performs an adjustment process for the indirect sound component of the second speaker. The sound processing device according to claim 13.
15. the position detection unit determines position information of the first speaker based on the sound signal. The sound processing device according to any one of claims 10 to 14.
16. It also has a camera, the camera acquires a video signal including facial images of the first speaker and the second speaker; the position detection unit determines position information of the first speaker based on the video signal. The sound processing device according to any one of claims 10 to 14.
17. the camera performs framing processing of the first speaker based on position information of the first speaker. The sound processing device according to claim 16.
18. Further comprising a sound image localization processing unit, the sound image localization processing unit performs localization processing to localize the direct sound component of the first speaker based on the position information of the first speaker. The sound processing device according to any one of claims 10 to 14.
Citation Information
Patent Citations
Information processing device, information processing terminal, information processing method, and program
WO2022054900A1