Information processing method, information processing device, and information processing program

By pitch-shifting and delaying sound data synthesis, the method addresses discomfort caused by dissonant intervals, enhancing sound accessibility for hearing-impaired individuals.

WO2026004741A1PCT designated stage Publication Date: 2026-01-02PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/022108
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-28
Filing Date
2025-06-19
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing sound processing technologies risk causing discomfort to listeners due to dissonant intervals when synthesizing audio data with different frequencies, particularly in environments where hearing-impaired individuals are present.

Method used

A method involving pitch-shifting sound data, setting a synthesis delay time based on time-series information, and delaying the second sound data to synthesize it with the first sound data, thereby reducing the likelihood of discomfort.

Benefits of technology

The method effectively generates synthetic sound that is less likely to cause discomfort, making it easier for hearing-impaired individuals to hear by adjusting frequency bands and reducing dissonant intervals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025022108_02012026_PF_FP_ABST
    Figure JP2025022108_02012026_PF_FP_ABST
Patent Text Reader

Abstract

This information processing method comprises: acquiring first sound data; generating second sound data obtained by pitch-shifting the first sound data; setting a synthesis delay time on the basis of time-series information of the first sound data; delaying the second sound data by the synthesis delay time and synthesizing the second sound data with the first sound data; and outputting third sound data obtained by the synthesis.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing method, information processing device, and information processing program

[0001] The present disclosure relates to processing sound data.

[0002] A known technology for processing sound data is described in Patent Document 1. Patent Document 1 discloses a voice guidance device that generates mixed voice data by combining pre-stored voice data of a guidance voice with voice data of a different frequency that is generated from the pre-stored voice data.

[0003] In the technology of Patent Document 1, depending on the combination of frequencies of the synthesized audio data, there is a risk that the listener (user) may feel uncomfortable when the mixed audio data is output.

[0004] Patent No. 4483450

[0005] The present disclosure has been made in consideration of the above circumstances, and aims to provide a technology that can generate synthetic sound in a manner that is less likely to cause discomfort to listeners.

[0006] In order to solve the above problem, a method according to one aspect of the present disclosure is an information processing method on a computer, including: acquiring first sound data; generating second sound data by pitch-shifting the first sound data; setting a synthesis delay time based on time-series information of the first sound data; delaying the second sound data by the synthesis delay time and synthesizing it with the first sound data; and outputting third sound data obtained by the synthesis.

[0007] According to the present disclosure, synthetic sound can be generated in a manner that is less likely to cause discomfort to listeners.

[0008] 1 is a block diagram illustrating a functional configuration of an information processing device according to an embodiment of the present disclosure. FIG. 2 is a flowchart illustrating a control operation by the information processing device. FIG. 3 is a subroutine illustrating details of the processing of step S3 of FIG. 2. FIG. 4 is a subroutine illustrating details of the processing of step S5 of FIG. 2. FIG. 5 is a subroutine illustrating details of the processing of step S6 of FIG. 2. FIG. 6 is a graph illustrating a sine wave of sound source data. FIG. 7 is a graph illustrating a spectrogram of acquired sound data. FIG. 8 is a graph illustrating time changes in the average power level of a target frequency band in the acquired sound data. FIG. 9 is a graph illustrating an example of the hearing characteristics of an elderly person. FIG. 10 is a graph for explaining a method for analyzing the average power level of the target frequency band in units of time frames. FIG. 11 is a graph for explaining a method for synthesizing pseudo natural sound data with the acquired sound data. FIG. 12 is a graph illustrating a case where the number of peaks (first feature), which is an example of an unpleasantness feature of synthetic sound data, is large. FIG. 13 is a graph illustrating a case where the number of peaks is small. FIG. 14 is a graph illustrating a case where power antagonism (second feature), which is another example of the unpleasantness feature, is large. FIG. 15 is a graph illustrating a case where the power antagonism is small. FIG. 16 is a graph illustrating a combination of peak frequencies of the synthetic sound data.

[0009] (Findings underlying the present disclosure) For example, when sound data is output from speakers or the like in facilities where hearing-impaired people such as elderly people or the hearing-impaired live, there is a need to output the sound data in a manner that is easy for hearing-impaired people to hear. One method that can meet this need is to perform processing (pitch shift) that changes the frequency band of the output sound data in consideration of the hearing characteristics of the hearing-impaired people.

[0010] As an example of a technology for changing the frequency band as described above, Patent Document 1 discloses a technology for processing guidance voices used in car navigation systems, elevators, etc. so that they are easier for users with hearing impairments to hear. Specifically, Patent Document 1 discloses generating two or more voice data with different frequencies from guidance voice voice data stored in a storage means such as a memory, synthesizing the generated voice data with original voice data, i.e., the voice data stored in the storage means, and outputting the resulting mixed voice data. Patent Document 1 considers that the mixed voice data includes data for each frequency band, such as high, low, and mid-range, and therefore the guidance voice obtained by outputting such mixed voice data will be easier for users with hearing impairments in some frequency bands to hear.

[0011] However, according to the knowledge of the present inventors, simply synthesizing a plurality of sound data of different frequency bands may generate a sound that causes discomfort to the listener (user), such as a sound including a dissonant interval. The inventors have found that synthesizing sound data with a time shift is effective in preventing such discomfort, and have arrived at the present disclosure.

[0012] (1) A method according to one aspect of the present disclosure is an information processing method on a computer, including: acquiring first sound data; generating second sound data by pitch-shifting the first sound data; setting a synthesis delay time based on time-series information of the first sound data; delaying the second sound data by the synthesis delay time and synthesizing the second sound data with the first sound data; and outputting third sound data obtained by the synthesis.

[0013] According to the present disclosure, acquired first sound data and second sound data obtained by pitch-shifting the first sound data are synthesized after being shifted in time by a synthesis delay time determined from the time series information of the first sound data, thereby preventing listeners from feeling uncomfortable when the synthesized third sound data is output. For example, if the first sound data and the second sound data are synthesized without a time shift, when the power level of the first sound data is relatively high, the second sound data obtained by pitch-shifting such first sound data is directly synthesized with the first sound data, which is likely to generate third sound data with large peaks in each frequency band before and after the pitch shift. Such third sound data may include unpleasant sounds, such as dissonant intervals. In contrast, according to the present disclosure, in which the first sound data and the second sound data are synthesized with a time shift, the above-mentioned phenomenon is less likely to occur, thereby preventing listeners from feeling uncomfortable.

[0014] (2) In the method (1) above, the second sound data may be generated by pitch-shifting data in a predetermined target frequency band of the first sound data.

[0015] In this aspect, by pitch-shifting data in a specific frequency band that is difficult for the hearing impaired to hear, it is possible to generate a synthesized sound (third sound data) that is easy for the hearing impaired to hear.

[0016] (3) In the method of (2) above, the second sound data may be generated by pitch-shifting data in the target frequency band of the first sound data to a lower pitch.

[0017] In this embodiment, it is possible to generate a synthetic sound (third sound data) suitable for hearing-impaired people such as elderly people who have difficulty hearing high-pitched sounds.

[0018] (4) In the method of (3) above, generating the second sound data may include calculating a centroid frequency of the target frequency band in the first sound data, setting a target frequency lower than the centroid frequency, calculating a pitch shift amount according to a difference between the centroid frequency and the target frequency, and pitch-shifting the data of the target frequency band by the calculated pitch shift amount.

[0019] In this embodiment, it is possible to determine an appropriate pitch shift amount according to the characteristics of a hearing-impaired person, and to make the pitch-shifted sound data easier for the hearing-impaired person to hear.

[0020] (5) In any of the methods (2) to (4) above, the synthesis delay time may be set based on a time from when the power level of the target frequency band in the first sound data rises above a predetermined value to when it falls below the predetermined value.

[0021] In this aspect, an appropriate synthesis delay time can be set according to the change over time in the power level of the frequency band to be pitch-shifted, thereby reducing the possibility that dissonant intervals will be included in the synthesized sound (third sound data).

[0022] (6) In any of the methods (1) to (5) above, the method may further include calculating a feature value representing an unpleasant feeling caused to a listener when the third sound data is output, and if the calculated feature value is equal to or greater than a predetermined threshold, resetting the synthesis delay time and regenerating the third sound data using the reset synthesis delay time.

[0023] In this embodiment, it is possible to accurately generate third sound data that does not cause discomfort to the listener.

[0024] (7) In the method of (6) above, the feature may be the number of peaks in the frequency characteristics of the third sound data, or a degree of power antagonism that indicates the proximity of power levels of each band on the low-pitched and high-pitched sides in the frequency characteristics.

[0025] In this aspect, it is possible to appropriately determine whether dissonant intervals or the like are likely to occur based on the characteristic amount.

[0026] (8) In any of the methods (1) to (5) above, the method may further include specifying a peak frequency or a scale of the generated third sound data, and determining whether or not a combination of the specified peak frequencies or scales approximates a predetermined unpleasant pattern; and, if it is determined that the combination approximates the unpleasant pattern, resetting the synthesis delay time and regenerating the third sound data using the reset synthesis delay time.

[0027] This method also makes it possible to accurately generate third sound data that does not cause discomfort to the listener.

[0028] (9) In any of the methods (1) to (8) above, the method may further include determining whether the acquired first sound data satisfies a predetermined condition indicating the need for pitch shifting, and the second sound data may be generated when the first sound data satisfies the predetermined condition.

[0029] In this aspect, in a situation where second sound data needs to be generated by pitch shifting, third sound data including the second sound data can be appropriately generated.

[0030] (10) In the method of (9) above, the predetermined condition may be that the first sound data includes a natural sound.

[0031] (11) In the method of (10) above, the predetermined condition may be that the first sound data includes the natural sound in a predetermined target frequency band.

[0032] In this aspect, by pitch-shifting data containing natural sounds in a predetermined frequency band, it is possible to generate second sound data (pseudo natural sound data) that is easy for hearing-impaired people to hear.

[0033] (12) A device according to another aspect of the present disclosure is an information processing device including a processor, the processor including: an acquisition unit that acquires first sound data; a generation unit that generates second sound data by pitch-shifting the first sound data; a setting unit that sets a synthesis delay time based on time series information of the first sound data; a synthesis unit that delays the second sound data by the synthesis delay time and synthesizes it with the first sound data; and an output unit that outputs third sound data obtained by the synthesis.

[0034] According to the present disclosure, it is possible to obtain the same effects as those of the information processing method disclosed above.

[0035] (13) According to yet another aspect of the present disclosure, there is provided an information processing program that causes a computer to acquire first sound data, generate second sound data by pitch-shifting the first sound data, set a synthesis delay time based on time series information of the first sound data, delay the second sound data by the synthesis delay time and then synthesize the second sound data with the first sound data, and output third sound data obtained by the synthesis.

[0036] The present disclosure can also provide the same effects as those of the information processing method disclosed above.

[0037] The present disclosure can also be realized as an information updating system that operates by such an information processing program. Needless to say, such a computer program can be distributed on a non-transitory computer-readable recording medium such as a CD-ROM or via a communication network such as the Internet.

[0038] Note that each of the embodiments described below represents a specific example of the present disclosure. The numerical values, shapes, components, steps, and step orders shown in the following embodiments are merely examples and are not intended to limit the present disclosure. Furthermore, among the components in the following embodiments, components that are not described in the independent claims that represent the highest concept are described as optional components. Furthermore, in all of the embodiments, the respective contents can be combined.

[0039] (Embodiment of the Present Disclosure) <Overall Configuration> Fig. 1 is a block diagram showing the functional configuration of an information processing device 1 according to an embodiment of the present disclosure. The information processing device 1 is a device that processes sound source data DS and outputs the processed data to a speaker 30, and is embodied by, for example, a microcomputer including a processor, a memory, a communication unit, and the like. The information processing device 1 functionally includes an acquisition unit 11, a natural sound extraction unit 12, a pseudo-natural sound generation unit 13, a delay time setting unit 14, a sound synthesis unit 15, a feature calculation unit 16, an discomfort determination unit 17, an output unit 18, and an equalizing unit 19. The pseudo-natural sound generation unit 13 corresponds to an example of a "generation unit" in the present disclosure, the delay time setting unit 14 corresponds to an example of a "setting unit" in the present disclosure, and the sound synthesis unit 15 corresponds to an example of a "synthesis unit" in the present disclosure.

[0040] The acquisition unit 11 is a module that acquires sound source data DS from the sound source memory 20. The sound source data DS may be data that is pre-stored in the sound source memory 20, or data that is temporarily stored in the sound source memory 20 via communication means from a mobile device such as a smartphone. The sound source data DS may include, for example, data on environmental sounds intended to provide a relaxing effect to the listener. As an example, the environmental sounds included in the sound source data DS may be a mixture of natural sounds such as the chirping of birds or insects or the flow of a river, and artificial sounds such as music.

[0041] The natural sound extraction unit 12 is a module that extracts natural sounds from the sound source data DS acquired by the acquisition unit 11. Specifically, the natural sound extraction unit 12 extracts data in a relatively high frequency band included in the sound source data DS as data in a target frequency band that may contain natural sounds. The natural sound extraction unit 12 then calculates the power level of the extracted target frequency band and determines that the sound source data DS contains natural sounds based on the time variation of the calculated power level.

[0042] The pseudo-natural sound generation unit 13 is a module that generates pseudo-natural sound data, which is pseudo-natural sound with different pitches (pitches), when natural sounds are included in the sound source data DS. Specifically, the pseudo-natural sound generation unit 13 generates pseudo-natural sound data by performing a pitch shift that shifts the pitch (pitch) of the data in the target frequency band extracted by the natural sound extraction unit 12 toward the lower end by a predetermined amount. In other words, the pseudo-natural sound data in this embodiment is an artificial natural sound obtained by shifting data in a relatively high frequency band (target frequency band) that includes natural sounds toward the lower end.

[0043] The sound synthesis unit 15 is a module that synthesizes the sound source data DS acquired by the acquisition unit 11 and the pseudo-natural sound data generated by the pseudo-natural sound generation unit 13. The synthesized sound data, which is the sound data after synthesis, is data in which the sound source data DS containing the original natural sound is mixed with the pseudo-natural sound data after pitch shifting.

[0044] The delay time setting unit 14 is a module that sets a synthesis delay time ΔT ( FIG. 11 ), which is a delay time when synthesizing pseudo-natural sound data. Specifically, the delay time setting unit 14 sets the synthesis delay time ΔT based on the original data before synthesis, that is, the time-series information of the sound source data DS acquired by the acquisition unit 11. The pseudo-natural sound data is delayed by the synthesis delay time ΔT set in this way and then synthesized with the original sound source data DS.

[0045] The feature calculation unit 16 is a module that calculates the discomfort feature Q ( FIG. 5 ) of the synthetic sound data generated by the sound synthesis unit 15. That is, based on the frequency characteristics of the synthetic sound data, the feature calculation unit 16 calculates, as the discomfort feature Q, a feature that indicates the discomfort that a listener will feel when the synthetic sound data is output. Note that the discomfort referred to here may be a clearly unpleasant sensation, or an unnatural sensation (uncomfortable feeling) that deviates from a pleasant sensation.

[0046] The discomfort determination unit 17 is a module that determines whether or not the listener feels uncomfortable based on the discomfort feature Q calculated by the feature calculation unit 16. Based on the above-described characteristics of the discomfort feature Q, the discomfort determination unit 17 determines that the output of the synthetic sound data will cause discomfort to the listener when the discomfort feature Q is equal to or greater than a predetermined threshold Qx ( FIG. 5 ).

[0047] The result of the determination by the discomfort determination unit 17 is fed back to the delay time setting unit 14. That is, when the discomfort determination unit 17 determines that discomfort has occurred, in other words, when the discomfort feature quantity Q is equal to or greater than the threshold value Qx, the delay time setting unit 14 resets the synthesis delay time ΔT. For example, the synthesis delay time ΔT is set so as to become longer each time the number of resets increases. When the synthesis delay time ΔT is reset, the sound synthesis unit 15 regenerates synthetic sound data using the reset synthesis delay time ΔT. This regeneration of synthetic sound data is repeated until the discomfort determination unit 17 determines that discomfort has not occurred (until the discomfort feature quantity Q becomes less than the threshold value Qx).

[0048] The output unit 18 is a module that outputs synthetic sound data that does not cause discomfort. That is, when the discomfort determination unit 17 determines that discomfort will not occur, that is, when the sound synthesis unit 15 obtains synthetic sound data in which the discomfort feature quantity Q is less than the threshold value Qx, the output unit 18 outputs the synthetic sound data to the equalizing unit 19.

[0049] The equalizing unit 19 is a module that performs equalization processing on the synthesized sound data output from the output unit 18. The equalizing processing is, for example, processing that finely adjusts frequency characteristics in accordance with the hardware characteristics of the speaker 30. The synthesized sound data that has been equalized by the equalizing unit 19 is output to the speaker 30. The speaker 30 reproduces the synthesized sound data and provides it to the listener. Note that the speaker 30 is not limited to a specific form as long as it has the function of reproducing the synthesized sound data and providing it to the listener. For example, the speaker 30 may be a stationary speaker that is installed in an appropriate location indoors or outdoors, or a wearable speaker that is worn on a part of the body like an earphone or a hearing aid.

[0050] <Control Operation> Next, a specific procedure for the control by the information processing device 1 as described above will be explained using the flowchart of Fig. 2. When the control shown in this figure starts, the acquisition unit 11 of the information processing device 1 sets 1 as the section number n of the sound source data DS (step S1). The section number n is the number of each section when the period from the start to the end of the sound source data DS is divided into fixed periods (for example, 10 seconds). Fig. 6 is a graph showing time-series data of the stereo signal of the sound source data DS, that is, a graph showing the change in amplitude (vertical axis) according to time (horizontal axis). As shown in this figure, in this embodiment, when the sound source data DS is read, the sound source data DS is divided into a plurality of sections C arranged in chronological order. 1 , C 2 ,...C N The section number setting (n=1) in step S1 is 1 , C 2 ,...C N This is a process of setting the number of the section to be read first to 1. 1 In the following, the sound data is read in order from each section C of the sound source data DS. 1 , C 2 ,...C N When referring to this without any particular distinction, it will be simply referred to as section C.

[0051] Next, the acquisition unit 11 acquires the n-th section C, which is data of the specified section C of the sound source data DS. n As described above, in step S1, the section number n is set to 1, so in step S2 immediately after step S1, the sound data of the first section C 1 After that, step S2 is repeated while the section number n is sequentially incremented through step S8 described later, and the sound data of each section C 1 , C 2 ,...C N In the following, the sound data of the nth section C n The sound data may be referred to as acquired sound data. The acquired sound data corresponds to an example of "first sound data" in the present disclosure.

[0052] Next, the natural sound extraction unit 12 of the information processing device 1 extracts the acquired sound data, that is, the nth section C n By analyzing the sound data, the section average level SL and standard deviation σ of the target frequency band Bt (FIG. 7) are calculated (step S3).

[0053] FIG. 3 is a subroutine showing details of the processing of step S3 above. As shown in this figure, when this processing starts, the natural sound extraction unit 12 performs frequency analysis of the acquired sound data (step S11). FIG. 7 shows a spectrogram obtained as a result of this frequency analysis. Specifically, the spectrogram shown in FIG. 7 is a graph showing changes in frequency over time, obtained by short-time Fourier analysis of the acquired sound data. The analysis of step S11 above is performed to obtain such frequency characteristics over time.

[0054] Next, the natural sound extraction unit 12 extracts data of a target frequency band Bt from the spectrogram (frequency characteristics) of the acquired sound data obtained as described above (step S12). The target frequency band Bt is set in advance as a relatively high frequency band that may include natural sounds, and in the example of Fig. 7, it is set to a band from a lower limit frequency f1 to an upper limit frequency f2. The frequencies f1 and f2 can be set in various ways, but for example, the lower limit frequency f1 may be set within the range of 2000 to 3000 Hz.

[0055] Next, the natural sound extraction unit 12 calculates the section average level SL shown in Fig. 8 as the time average of the power level (dB) of the extracted target frequency band Bt (step S13). The graph in Fig. 8 shows how the average power level La of the target frequency band Bt in the acquired sound data changes over time. The average power level La is a value obtained by averaging the power levels of each frequency (f1 to f2) included in the target frequency band Bt, and is calculated for each time period. From this time-series data of the average power level La of the target frequency band Bt, the natural sound extraction unit 12 calculates the section average level SL by further averaging the average power level La over time within the section C of the acquired sound data.

[0056] Next, the natural sound extraction unit 12 calculates the standard deviation σ, which represents the magnitude of time fluctuation in the average power level La of the target frequency band Bt (step S14). As shown in Fig. 8, the standard deviation σ is a value that statistically represents the magnitude of time fluctuation in the average power level La centered around the section average level SL, and the larger this value is, the greater the fluctuation of the average power level La within the section.

[0057] 2 and determines whether or not the acquired sound data includes natural sound (step S4). This determination is for generating pseudo-natural sound data (S5) described below, in other words, for checking the necessity of pitch-shifting the natural sound, and corresponds to an example of a determination as to whether or not the "predetermined condition indicating the necessity of pitch-shifting" in the present disclosure is satisfied.

[0058] Specifically, in step S4 above, the natural sound extraction unit 12 determines whether the following conditions are met: the section average level SL is greater than a predetermined reference value SLx, and the standard deviation σ is greater than a predetermined reference value σx. If these conditions are met, it is determined that the acquired sound data contains natural sounds. Natural sounds such as the chirping of birds or insects or the flow of a river generally have large temporal fluctuations. Therefore, if a certain amount of natural sounds are included in the acquired sound data, both the section average level SL and the standard deviation σ of the target frequency band Bt should be large. The above determination is made with these circumstances in mind.

[0059] If the determination in step S4 is NO, that is, if at least one of the section average level SL and the standard deviation σ is equal to or less than the reference value (SLx, σx) and the acquired sound data does not contain natural sound, the flow proceeds to step S7, which will be described later. In this case, the processes of generating and synthesizing pseudo-natural sound data (S5, S6), which will be described later, are not performed.

[0060] On the other hand, if the determination in step S4 is YES, that is, if it is confirmed that the section average level SL and the standard deviation σ are both greater than the reference values ​​(SLx, σx) and that the acquired sound data contains natural sounds, the pseudo-natural sound generation unit 13 of the information processing device 1 generates pseudo-natural sound data by pitch-shifting the data in the target frequency band Bt to the lower frequency side (step S5). Note that the pseudo-natural sound data corresponds to an example of the "second sound data" in the present disclosure.

[0061] 4 is a subroutine showing details of the processing of step S5 above. As shown in the figure, when this processing starts, the pseudo-natural sound generation unit 13 calculates the section centroid frequency fg of the target frequency band Bt (step S21). The section centroid frequency fg is a value obtained by further averaging the centroid frequencies at each time point of the target frequency band Bt over time. The centroid frequency is the average value of the frequencies of the target frequency band Bt weighted by the power level, and is calculated for each time period. From these centroid frequencies at each time, the pseudo-natural sound generation unit 13 calculates the section centroid frequency fg by further averaging the centroid frequencies over time within the section C of the acquired sound data.

[0062] Next, the pseudo-natural sound generator 13 determines the target frequency ft (step S22). The target frequency ft can be determined in advance based on the hearing characteristics of the assumed listener (hearing-impaired person). For example, if the assumed listener is an elderly person who has difficulty hearing high-pitched sounds, the target frequency ft can be determined in advance based on the hearing characteristics of this elderly person.

[0063] FIG. 9 shows the average hearing level Ha of a 70-year-old male as an example of the hearing characteristics of the elderly. Based on the hearing characteristics of the elderly, the target frequency ft can be set to an appropriate frequency that minimizes the difference in hearing ability between the elderly and those with normal hearing. Specifically, when the difference between the upper limit sound pressure level Z (dBHL), which is predetermined as the sound pressure level at which a person with normal hearing perceives a sound as being noisy, and the average hearing level Ha of the elderly is ΔH, the frequency at which this ΔH becomes a predetermined value can be set as the target frequency ft. By setting the target frequency ft in this manner, the output sound after pitch shifting can be made easy to hear for the elderly (hearing-impaired) while preventing people with normal hearing from perceiving it as being noisy.

[0064] Another method is to determine the target frequency ft by taking into account the listener's compression characteristics. The compression characteristics are those in which the increase in output sound pressure level decreases as the sound pressure level input to the cochlea increases, and are defined as the relationship between the input sound pressure level at the cochlear entrance and the output sound pressure level. Such compression characteristics vary with frequency, and it is known that the rate of increase in the mid-range sound pressure range is greater for hearing-impaired individuals than for those with normal hearing. Therefore, for example, the target frequency ft can be set to a frequency at which the rate of increase in a specific input level range (e.g., an input sound pressure level at the cochlear entrance of 40 to 60 dB) for a hearing-impaired individual, such as an elderly person, exceeds a predetermined value. In other words, the target frequency ft can be set to a frequency at which the slope of the graph, with the horizontal axis representing the input sound pressure level (dB) and the vertical axis representing the output sound pressure level (dB), is equal to or less than a predetermined slope. This allows the pitch-shifted frequency to be perceived as a frequency with compression characteristics similar to those perceived by those with normal hearing.

[0065] Once the target frequency ft has been determined in the above manner, the pseudo-natural sound generator 13 then calculates the pitch shift amount Ps (step S23). The pitch shift amount Ps is calculated using the following equation (1) based on the section centroid frequency fg calculated in step S21 and the target frequency ft determined in step S22.

[0066] Ps = 12 × log 2 (fg / ft) Formula (1)

[0067] The pitch shift amount Ps calculated by the above formula (1) is the difference between the section centroid frequency fg and the target frequency ft, expressed in semitones. For example, if the section centroid frequency fg is 4000 Hz and the target frequency ft is 2000 Hz, the pitch shift amount Ps is calculated to be 12 semitones.

[0068] Next, the pseudo-natural sound generation unit 13 shifts the pitch of the data of the target frequency band Bt toward the lower frequency side by the pitch shift amount Ps calculated as described above (step S24), and generates this pitch-shifted data as pseudo-natural sound data (step S25). The pseudo-natural sound data obtained in this manner is data in which the data of the target frequency band Bt has been shifted overall toward the lower frequency side so that the section center of gravity frequency fg of that band moves to the target frequency ft, in other words, data in which the pitch (pitch) of the target frequency band Bt has been lowered overall by the pitch shift amount Ps.

[0069] Once the pseudo-natural sound data has been generated in the above manner, the delay time setting unit 14 and sound synthesis unit 15 of the information processing device 1 return to the flow of Figure 2 and perform a process of synthesizing the generated pseudo-natural sound data with the acquired sound data (step S6).

[0070] FIG. 5 is a subroutine showing details of step S6. As shown in this figure, when this process starts, the delay time setting unit 14 of the information processing device 1 acquires the average power level La of the target frequency band Bt for each time frame F shown in FIG. 10 (step S31). The time frame F refers to each sub-section when the section C (FIG. 6) of the acquired sound data is divided into small periods (e.g., 20 msec). The delay time setting unit 14 acquires the power level of each time frame F by examining the change over time in the average power level La of the target frequency band Bt in the acquired sound data. The power level of each time frame F may be, for example, the value of the average power level La at the start or end of each time frame F, or may be the average value of the average power level La within each time frame F.

[0071] Next, the delay time setting unit 14 identifies, as the rising time frame Fu, the time frame F in which the power level first becomes equal to or greater than a predetermined value Y among all time frames F in section C of the acquired sound data (step S32). The predetermined value Y is set in advance to be a value between the maximum power level and the minimum power level in section C. In the example of Fig. 10, when the numbers of the time frames F are counted sequentially from the start of section C, the power level first becomes equal to or greater than the predetermined value Y in the time frame F with number K1. Therefore, the delay time setting unit 14 identifies the time frame F with number K1 as the rising time frame Fu.

[0072] Next, the delay time setting unit 14 identifies, as a fall time frame group Fd, a group of time frames F after the rise time frame Fu whose power levels are less than a predetermined value Y (step S33). In the example of FIG. 10 , the power levels of the time frames F with numbers K2, K3, K4, ..., which are greater than number K1 of the rise time frame Fu, are all less than the predetermined value Y. Therefore, the delay time setting unit 14 identifies the group of time frames F with numbers K2, K3, K4, ..., as a fall time frame group Fd. In other words, the fall time frame group Fd includes a first fall time frame Fd1 with number K2, a second fall time frame Fd2 with number K3, a third fall time frame Fd3 with number K4, ...

[0073] Next, the delay time setting unit 14 sets the first time of the fall time frame group Fd as the combined delay time ΔT (step S34). The first time of the fall time frame group Fd is the time of the first fall time frame Fd1, which has the smallest number in the fall time frame group Fd, relative to the rise time frame Fu. That is, the delay time setting unit 14 calculates the elapsed time ΔT1 from the rise time frame Fu to the first fall time frame Fd1 and sets this as the combined delay time ΔT. In other words, the combined delay time ΔT is set based on the time from when the power level of the target frequency band Bt rises above a predetermined value Y to when it falls below the predetermined value Y.

[0074] Next, the sound synthesis unit 15 of the information processing device 1 synthesizes the acquired sound data and the pseudo natural sound data (step S35). That is, the sound synthesis unit 15 delays the pseudo natural sound data generated in step S25 by the synthesis delay time ΔT, and then synthesizes the pseudo natural sound data for the n-th interval C read in step S2. n 11 , when the acquired sound data is D1 and the pseudo natural sound data is D2, the sound synthesis unit 15 shifts the pseudo natural sound data D2 to the delay side (right side) on the time axis so that output of the pseudo natural sound data D2 starts with a delay of synthesis delay time ΔT from the start of the acquired sound data D1. Then, by synthesizing the shifted pseudo natural sound data D2 with the acquired sound data D1, synthetic sound data including both data D1 and D2 is generated. Note that the synthetic sound data corresponds to an example of the "third sound data" in the present disclosure.

[0075] Next, the feature calculation unit 16 of the information processing device 1 calculates the discomfort feature Q based on the frequency characteristics of the synthetic sound data generated in step S35 (step S36). As described above, the discomfort feature Q is a feature representing the discomfort that a listener is likely to experience when synthetic sound data is output. In this embodiment, the discomfort feature Q is predefined as an amount that increases in a situation where a pseudo-natural sound and an original natural sound are simultaneously output at a high volume. A large discomfort feature Q indicates that dissonance and the like are likely to occur, and that the listener is likely to experience discomfort.

[0076] Specific examples of the discomfort feature quantity Q include the following first feature quantity Q1 and second feature quantity Q2.

[0077] The first feature Q1 is the number of peaks exceeding a predetermined reference value, i.e., the number of peaks. For example, when an original natural sound is loud and a pseudo-natural sound pitch-shifted to the lower end is simultaneously output, peaks with high power levels occur in both the low and high frequency ranges, resulting in a natural increase in the number of peaks exceeding the reference value. Conversely, when the original natural sound and the pseudo-natural sound are not simultaneously output, the number of peaks exceeding the reference value is reduced. In other words, when the number of peaks exceeding the reference value is large, dissonant intervals, which are unpleasant sound combinations, are more likely to occur, while when the number of peaks exceeding the reference value is small, dissonant intervals are less likely to occur. Therefore, in this case, as shown in Figures 12A and 12B, the number of peaks whose peak values ​​(maximum power levels) exceed a predetermined reference value V among the peaks appearing in the graph of the frequency characteristics of the synthetic sound data is counted, and this number of peaks is used as the first feature Q1. The larger the first feature Q1 (number of peaks), the more likely it is that the listener will experience discomfort when the synthetic sound data is output. Note that Figure 12A represents a situation in which the number of peaks is large and the first feature Q1 is large, i.e., a situation in which discomfort is likely to occur, while Figure 12B represents a situation in which the number of peaks is small and the first feature Q1 is small, i.e., a situation in which discomfort is unlikely to occur.

[0078] The second feature Q2 is the degree of power antagonism, which represents the closeness of the power levels in the low and high frequency ranges. For example, when an original natural sound and a pseudo-natural sound are output simultaneously, the power level in the high frequency range (high frequency range) including the original natural sound and the power level in the low frequency range (low frequency range) including the pseudo-natural sound become closer, resulting in a larger degree of power antagonism. Conversely, when an original natural sound and a pseudo-natural sound are not output simultaneously, the degree of power antagonism becomes smaller. In other words, when the degree of power antagonism is large, dissonant intervals and the like are more likely to occur, and when the degree of power antagonism is small, dissonant intervals and the like are less likely to occur. Therefore, in this case, as shown in Figures 13A and 13B, the average power level L1 of the high frequency range B1 and the average power level L2 of the low frequency range B2 are examined, and the second feature Q2, which represents the closeness of the two power levels L1 and L2, is calculated using the following formula (2): 13A and 13B, the treble range B1 is the same as or includes the target frequency band Bt (FIG. 7), and the bass range B2 is a band with a lower frequency than the treble range B1. The treble range B1 and the bass range B2 are adjacent to each other, and the boundary between them can be, for example, the same as the lower limit frequency f1 of the target frequency band Bt. In other words, the treble range B1 is the band before pitch shifting, and the bass range B2 is the band after pitch shifting.

[0079] Q2=min(L1,L2) / max(L1,L2)...Formula (2)

[0080] As shown in equation (2), the second feature Q2 is the ratio between the average power level L1 of the treble range B1 and the average power level L2 of the bass range B2. More specifically, the second feature Q2 is a value of 1 or less obtained by dividing the smaller of the two average power levels L1 and L2 by the larger of the two average power levels L1 and L2. The larger this second feature Q2 (power antagonism) is (the closer it is to 1), the closer the power levels L1 and L2 of the two ranges B1 and B2 are, and the more likely it is that the listener will experience discomfort when the synthesized sound data is output. Note that Figure 13A illustrates a situation in which the power levels of the treble range B1 and the bass range B2 are close to each other, resulting in a large second feature Q2 (power antagonism), i.e., a situation in which discomfort is likely to occur. Meanwhile, Figure 13B illustrates a situation in which the power levels of the treble range B1 and the bass range B2 are far apart, resulting in a small second feature Q2 (power antagonism), i.e., a situation in which discomfort is unlikely to occur.

[0081] Here, either the first feature quantity Q1 or the second feature quantity Q2 described above may be used as the discomfort feature quantity Q, or a combination of both feature quantities Q1 and Q2 may be used. In the latter case, one of the first and second feature quantities Q1 and Q2 may be calculated first to make a determination (S37) described later, and depending on the result of that determination, the other of the first and second feature quantities Q1 and Q2 may then be calculated and made a determination, so that calculation and determination of the feature quantities may be performed in stages.

[0082] When the discomfort feature Q (e.g., the first feature Q1 or the second feature Q) is calculated in step S36, the discomfort determination unit 17 of the information processing device 1 determines whether the calculated discomfort feature Q is less than a predetermined threshold Qx (step S37).

[0083] The threshold Qx used in the determination of step S37 above can be set to, for example, 3 when the discomfort feature Q is the above-mentioned first feature Q1, i.e., the number of peaks. When the discomfort feature Q is the above-mentioned second feature Q2, i.e., the power antagonism, the threshold Qx can be set to, for example, 0.5. Note that the threshold Qx for the second feature Q2 (power antagonism) is not limited to 0.5 and can be set to an appropriate value within the range of, for example, 0.4 to 0.6.

[0084] Furthermore, since the number of peaks or the degree of power antagonism described above are values ​​found from the frequency characteristics of each time in the synthetic sound data, the discomfort feature Q can be calculated for each time frame F shown in FIG. 10 . Therefore, the determination in step S37 described above may employ the following determination method that takes this into consideration. For example, the condition for a NO determination in step S37 may be that the discomfort feature Q is equal to or greater than the threshold Qx in at least one time frame F. Alternatively, the condition for a NO determination in step S37 may be that the discomfort feature Q is equal to or greater than the threshold Qx in a specific number or more time frames F that are greater than 1.

[0085] If the determination in step S37 is YES and it is confirmed that the discomfort feature quantity Q is less than the threshold value Qx, the sound synthesis unit 15 applies the synthetic sound data obtained in the immediately preceding step S35 to the currently targeted section C (the n-th section C n ) is confirmed as the synthesized voice data (step S38).

[0086] On the other hand, if the determination in step S37 is NO and it is confirmed that the discomfort feature quantity Q is equal to or greater than the threshold Qx, the delay time setting unit 14 resets the synthesis delay time ΔT to the next time in the fall time frame group Fd ( FIG. 10 ) described above (step S39). For example, if the synthesis delay time ΔT used when generating the immediately preceding synthetic sound data was the time corresponding to the first fall time frame Fd1, the delay time setting unit 14 resets the synthesis delay time ΔT to the time corresponding to the next second fall time frame Fd2, i.e., the elapsed time ΔT2 from the rise time frame Fu to the second fall time frame Fd2. Similarly, if the synthesis delay time ΔT used immediately before was the time corresponding to the second fall time frame Fd2, the delay time setting unit 14 resets the synthesis delay time ΔT to the time corresponding to the next third fall time frame Fd3, i.e., the elapsed time ΔT3 from the rise time frame Fu to the third fall time frame Fd3. In this way, the combined delay time ΔT is increased every time it is determined that the discomfort feature quantity Q is equal to or greater than the threshold value Qx.

[0087] The synthesis delay time ΔT reset as described above is used again to generate synthetic sound data. That is, the sound synthesis unit 15 delays the pseudo-natural sound data by the reset synthesis delay time ΔT and then synthesizes it with the acquired sound data (step S35). In this way, the synthetic sound data is regenerated.

[0088] Thereafter, the discomfort feature Q of the regenerated synthetic sound data is calculated (step S36), and a determination is made regarding the discomfort feature Q (step S37). The synthesis delay time ΔT is reset (S39) and the synthetic sound data is regenerated (S35) repeatedly until the determination result becomes YES, that is, until the discomfort feature Q becomes less than the threshold Qx. The synthetic sound data finally obtained by this repetition is the target section C (nth section C n ) is determined as the synthesized speech data (step S38).

[0089] 2, and determines whether acquisition of sound data for all sections C of the sound source data DS in step S2 has been completed (step S7). For example, as shown in FIG. 6, if there are a total of N sections C of the sound source data DS, the sound synthesis unit 15 determines whether the number (n) of the section C from which sound data was last acquired matches N. If they match (n=N), the sound synthesis unit 15 determines that sound data for all sections C has been acquired.

[0090] If the determination in step S7 is NO and it is confirmed that there remains a section C for which sound data has not been acquired, the sound synthesis unit 15 sets the section number of the sound source data DS to be read next by the acquisition unit 11 to n+1, which is 1 greater than the current number (step S8). n+1) sound data is acquired by the acquisition unit 11, and the acquired sound data is subjected to the processing from step S3 onwards. This processing is repeated until sound data for the entire section C is acquired. As a result, it is determined whether natural sounds are included in the entire section C of the sound source data DS, and synthetic sound data including pseudo-natural sound data is generated for all sections C that include natural sounds.

[0091] Next, the sound synthesis unit 15 combines the sound data for all intervals C to generate output sound data (step S9). The output sound data is transmitted to the speaker 30 via the output unit 18 and equalizing unit 19 shown in FIG. 1. If natural sounds are included in all intervals C of the sound source data DS, pseudo-synthetic sounds are generated and synthesized (S5, S6) for all intervals C, and therefore the output sound data transmitted to the speaker 30 is entirely composed of synthetic sound data including pseudo-natural sound data. On the other hand, if natural sounds are included only in a portion of interval C, the output sound data transmitted to the speaker 30 is a combination of synthetic sound data including pseudo-natural sound data and the original sound data (acquired sound data).

[0092] <Effects> As described above, in this embodiment, acquired sound data acquired from the sound source memory 20 and pseudo-natural sound data obtained by pitch-shifting data in the target frequency band Bt of the acquired sound data to the lower frequency side are synthesized after being shifted in time by the synthesis delay time ΔT determined from the time series information of the acquired sound data ( FIG. 10 ). This makes it possible to prevent listeners from feeling uncomfortable when the synthesized synthetic sound data (or output sound data including the synthesized synthetic sound data) is output. For example, if the acquired sound data and pseudo-natural sound data are synthesized without a time shift, when the power level of the target frequency band Bt of the acquired sound data is relatively high, the pseudo-natural sound data obtained by pitch-shifting the data in the target frequency band Bt to the lower frequency side is directly synthesized with the acquired sound data, which tends to generate synthetic sound data with large peaks in each frequency band (high and low frequency ranges) before and after the pitch shift. Such synthetic sound data may include unpleasant sounds, such as dissonant intervals. In contrast, according to the present embodiment, in which the acquired sound data and the pseudo natural sound data are synthesized with a time lag, the above-mentioned phenomenon is less likely to occur, and therefore discomfort felt by the listener can be suppressed.

[0093] In particular, in this embodiment, a relatively high frequency band that can contain natural sounds is set as the target frequency band Bt, and pseudo-natural sound data is generated by pitch-shifting the data of this target frequency band Bt toward the lower frequency side.Therefore, by synthesizing this pseudo-natural sound data with the original acquired sound data, synthetic sound data that is suitable for hearing-impaired people such as the elderly who have difficulty hearing high-pitched sounds can be generated.

[0094] Furthermore, in this embodiment, the pitch shift amount Ps when pitch-shifting the data of the target frequency band Bt is calculated according to the difference between the section centroid frequency fg (centroid frequency) of the target frequency band Bt and a target frequency ft that is predetermined to be lower than the section centroid frequency fg. With this configuration, it is possible to determine an appropriate pitch shift amount Ps according to the characteristics of a hearing-impaired person, and to make the pseudo-natural sound data after pitch shift easier to hear for a hearing-impaired person.

[0095] Furthermore, in this embodiment, the synthesis delay time ΔT is set based on the time from when the power level of the target frequency band Bt rises above a predetermined value Y to when it falls below the predetermined value Y. With this configuration, an appropriate synthesis delay time ΔT can be set in accordance with the change over time in the power level of the target frequency band Bt, thereby reducing the possibility that dissonant intervals will be included in the synthesized sound data.

[0096] Furthermore, in this embodiment, an discomfort feature Q that indicates the discomfort felt by the listener when synthetic sound data is output is calculated, and if the discomfort feature Q is equal to or greater than a predetermined threshold Qx, the synthesis delay time ΔT is reset. Then, synthetic sound data is regenerated using the reset synthesis delay time ΔT. With this configuration, the synthesis delay time ΔT can be repeatedly adjusted until the discomfort feature Q becomes less than the threshold Qx, and synthetic sound data that does not cause discomfort to the listener can be accurately generated.

[0097] Specifically, at least one of a first feature Q1, which is the number of peaks exceeding a predetermined reference value (peak count), and a second feature Q2, which is the power antagonism that indicates the proximity of the power levels of the low and high frequency ranges, can be used as the unpleasantness feature Q. By using such features, it is possible to appropriately determine whether dissonant intervals or the like are likely to occur.

[0098] In addition, in this embodiment, it is determined in advance whether or not natural sounds are included in the target frequency band Bt of the acquired sound data, and if natural sounds are included, pseudo-natural sound data is generated by pitch-shifting the data of the target frequency band Bt to the lower frequency side. With this configuration, by pitch-shifting the data of a relatively high frequency band that includes natural sounds, pseudo-natural sound data that is easy for hearing-impaired people such as elderly people to hear can be generated.

[0099] <Modifications> Although the preferred embodiments of the present disclosure have been described above, the present disclosure is not limited to these, and the following modifications can be adopted, for example.

[0100] In the above embodiment, the synthesis delay time ΔT when synthesizing the acquired sound data (first sound data) with the pseudo-natural sound data (second sound data) is first set, and then the discomfort feature Q such as the number of peaks and the degree of power antagonism is calculated from the frequency characteristics of the synthesized sound data (third sound data), which is the data after synthesis, and the synthesis delay time ΔT is reset if the calculated discomfort feature Q is equal to or greater than a predetermined threshold value Qx. However, the method of determining whether to reset the synthesis delay time ΔT is not limited to this.

[0101] For example, as shown in FIG. 14 , peak frequencies fp1, fp2, ..., which are frequencies of peaks whose power levels exceed a predetermined value, may be identified from the frequency characteristics of the synthetic sound data (third sound data), and a determination may be made as to whether the combination of the identified peak frequencies fp1, fp2, ..., approximates a predetermined discomfort pattern. If so, the synthesis delay time ΔT may be reset. The discomfort pattern may be, for example, a combination of a predetermined first frequency and a predetermined second frequency. In this case, if there is a combination of frequencies approximate to the first frequency and the second frequency among the peak frequencies fp1, fp2, ..., of the synthetic sound data, the synthesis delay time ΔT may be reset. The first frequency may be set to, for example, 2500 Hz, and the second frequency may be set to, for example, 4500 Hz. Furthermore, whether or not the peak frequencies fp1, fp2, ... are similar to the first and second frequencies, i.e., whether or not the peak frequencies are similar to the discomfort pattern, can be determined by checking whether or not there is a combination of frequencies among the peak frequencies fp1, fp2, ... whose difference from the first and second frequencies is equal to or smaller than a predetermined allowable error (for example, ±200 Hz).

[0102] Furthermore, instead of examining a combination of frequencies as described above, a combination of musical scales may be examined. That is, the musical scale of the synthesized sound data (third sound data) may be identified, and if the identified musical scale resembles a predetermined unpleasant pattern, the synthesis delay time ΔT may be reset. In this case, the unpleasant pattern may be, for example, a combination of musical scales that form a predetermined dissonant interval.

[0103] In the above embodiment, the ratio between the average power level L1 of the treble range B1 and the average power level L2 of the bass range B2 (the value obtained by dividing the smaller power level by the larger power level) is used as the power antagonism (second feature Q2), which is an example of the discomfort feature Q. However, the power antagonism is not limited to the above definition, and may be any value as long as it increases as the average power level L1 of the treble range B1 and the average power level L2 of the bass range B2 become closer. For example, the second feature Q2' defined by the following formula (3) may be used as the power antagonism. In other words, the power antagonism in this case is the reciprocal of the value obtained by dividing the absolute value of the difference between the average power level L1 of the treble range B1 and the average power level L2 of the bass range B2 by the larger power level (L1 or L2).

[0104] Q2'={|L2-L1| / max(L1,L2)} -1 ...Formula (3)

[0105] In the above embodiment, the section average level SL and standard deviation σ of the target frequency band Bt are calculated by examining the time change ( FIG. 8 ) of the average power level La of the target frequency band Bt in the acquired sound data (first sound data), and whether or not the acquired sound data contains natural sound is determined based on the calculated values ​​SL and σ. However, the method of determining whether or not natural sound is present is not limited to this. For example, the similarity between the time change of the average power level La of the target frequency band Bt and pre-registered natural sound data may be examined, and the presence or absence of natural sound may be determined from the result. In this case, the similarity with the natural sound data can be examined, for example, as follows. That is, the envelope of the waveform of the average power level La is identified, and the correlation coefficient between the identified envelope and the envelope of the registered waveform is used as the similarity.

[0106] In the above embodiment, the sound source data DS is divided into sections C 1 , C 2 ,...C N In the above example, sound data is read in for each sound data item, and processing such as pitch shifting is performed on each read sound data item (acquired sound data) as necessary. However, it is also possible to read in the entire sound source data DS all at once and then perform processing such as pitch shifting.

[0107] INDUSTRIAL APPLICABILITY The present disclosure can generate synthetic sound in a manner that is unlikely to cause discomfort to listeners, and is therefore useful for processing sound data to generate synthetic sound that is easy for listeners to hear.

Claims

1. An information processing method on a computer, comprising: acquiring first sound data; generating second sound data by pitch-shifting the first sound data; setting a synthesis delay time based on time series information of the first sound data; synthesizing the second sound data with the first sound data after delaying it by the synthesis delay time; and outputting third sound data obtained by the synthesis.

2. An information processing method according to claim 1, wherein the second sound data is generated by pitch-shifting data in a predetermined target frequency band of the first sound data.

3. An information processing method according to claim 2, wherein the second sound data is generated by pitch-shifting the data in the target frequency band of the first sound data to a lower pitch.

4. An information processing method according to claim 3, wherein generating the second sound data includes: calculating a centroid frequency of the target frequency band in the first sound data; setting a target frequency lower than the centroid frequency; calculating a pitch shift amount according to the difference between the centroid frequency and the target frequency; and pitch-shifting the data of the target frequency band by the calculated pitch shift amount.

5. An information processing method according to claim 2, wherein the synthesis delay time is set based on the time from when the power level of the target frequency band in the first sound data rises above a predetermined value to when it falls below the predetermined value.

6. An information processing method according to any one of claims 1 to 5, further comprising: calculating a feature representing the discomfort experienced by a listener when the third sound data is output; and, if the calculated feature is equal to or greater than a predetermined threshold, resetting the synthesis delay time and regenerating the third sound data using the reset synthesis delay time.

7. An information processing method according to claim 6, wherein the feature is the number of peaks in the frequency characteristics of the third sound data, or a degree of power antagonism that indicates the closeness of the power levels of the low-pitched and high-pitched bands in the frequency characteristics.

8. An information processing method according to any one of claims 1 to 5, further comprising: identifying peak frequencies or musical scales of the generated third sound data, and determining whether the combination of the identified peak frequencies or musical scales approximates a predetermined unpleasant pattern; and if it is determined that the combination approximates the unpleasant pattern, resetting the synthesis delay time, and regenerating the third sound data using the reset synthesis delay time.

9. An information processing method according to any one of claims 1 to 5, further comprising determining whether the acquired first sound data satisfies a predetermined condition indicating the need for pitch shifting, and wherein the second sound data is generated when the first sound data satisfies the predetermined condition.

10. An information processing method according to claim 9, wherein the predetermined condition is that the first sound data includes a natural sound.

11. An information processing method according to claim 10, wherein the predetermined condition is that the first sound data contains the natural sound in a predetermined target frequency band.

12. An information processing device including a processor, wherein the processor includes: an acquisition unit that acquires first sound data; a generation unit that generates second sound data by pitch-shifting the first sound data; a setting unit that sets a synthesis delay time based on time series information of the first sound data; a synthesis unit that delays the second sound data by the synthesis delay time and synthesizes it with the first sound data; and an output unit that outputs third sound data obtained by the synthesis.

13. An information processing program that causes a computer to execute the following steps: acquiring first sound data; generating second sound data by pitch-shifting the first sound data; setting a synthesis delay time based on time series information of the first sound data; synthesizing the second sound data with the first sound data after delaying it by the synthesis delay time; and outputting third sound data obtained by the synthesis.

Citation Information

Patent Citations

  • Device and method for voice guidance, and navigation device

    JP2006038929A

  • Navigation device, navigation device control method, program for the navigation device control method, and recoding medium with the program for navigation device control method stored thereon

    JP2007333603A