Sound reproduction method, recording medium, and sound reproduction system
By acquiring the user's head movement velocity and merging the convolution processing of the head-related transfer function, the problem of excessive stereo computation load is solved, achieving more appropriate computation processing and a better stereo experience.
Patent Information
- Application Number
- CN202180019555.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-12-17
- Filing Date
- 2021-03-04
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2041-03-04
AI Technical Summary
Existing technologies place an excessive computational load on stereo sound, leading to increased computing resources and power consumption, and limiting the user's stereo experience.
By acquiring the user's head movement speed, an output sound signal is generated, enabling the user to perceive sound arriving from different positions in the three-dimensional sound field. The convolution processing of the head-related transfer function is combined to reduce the amount of computation and suppress the perception of changes in the position of the sound image.
It reduces the computational load, decreases computing resources and power consumption, while improving the stereo user experience and avoiding the disharmony of sound image positioning.
Smart Images

Figure CN115244947B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to sound reproduction systems, programs, and methods. Background Technology
[0002] Previously, there were known techniques for enabling users to perceive stereo sound reproduction by controlling the position of the sound image of a perceptual sound source object within a virtual three-dimensional space (for example, see Patent Document 1).
[0003] Existing technical documents
[0004] Patent documents
[0005] Patent Document 1: Japanese Patent Application Publication No. 2020-18620 Summary of the Invention
[0006] The problem that the invention aims to solve
[0007] On the other hand, generating stereo sound for users to perceive requires massive computational processing. In contrast, conventional sound reproduction methods have sometimes lacked adequate computational processing.
[0008] In view of the above, the purpose of this disclosure is to provide a method for sound reproduction that enables users to perceive stereo sound through more appropriate computational processing.
[0009] Methods used to solve problems
[0010] The present disclosure discloses a sound reproduction method that enables a user to perceive a first sound as arriving from a first position in a three-dimensional sound field, and enables the user to perceive a second sound as arriving from a second position different from the first position. The method includes: an acquisition step of acquiring the movement speed of the user's head; and a generation step of generating an output sound signal for enabling the user to perceive a sound arriving from a predetermined position in the three-dimensional sound field. In the generation step, if the acquired movement speed is greater than a first threshold, the output sound signal is generated to enable the user to perceive both the first and second sounds as arriving from a third position between the first and second positions.
[0011] Furthermore, the sound reproduction system of one of the technical solutions disclosed herein enables a user to perceive a first sound as a sound arriving from a first position in a three-dimensional sound field, and enables the user to perceive a second sound as a sound arriving from a second position different from the first position. This system includes: an acquisition unit that acquires the movement speed of the user's head; and a generation unit that generates an output sound signal for enabling the user to perceive a sound arriving from a predetermined position in the three-dimensional sound field. If the acquired movement speed is greater than a first threshold, the generation unit generates the output sound signal for enabling the user to perceive both the first and second sounds as sounds arriving from a third position between the first and second positions.
[0012] Furthermore, the technical solution disclosed herein can also be implemented as a non-temporary recording medium readable by a computer that records a program for enabling a computer to execute the sound reproduction method described above.
[0013] In addition, these inclusive or specific technical solutions can also be implemented by non-transitory recording media such as systems, devices, methods, integrated circuits, computer programs, or computer-readable CD-ROMs, or by any combination of systems, devices, methods, integrated circuits, computer programs, and recording media.
[0014] Invention Effects
[0015] According to this disclosure, it is possible to make users perceive stereo sound through more appropriate computational processing. Attached Figure Description
[0016] Figure 1 This is a schematic diagram illustrating an example of the use of the sound reproduction system according to the implementation method.
[0017] Figure 2 This is a block diagram illustrating the functional structure of the sound reproduction system according to the implementation method.
[0018] Figure 3 This is a flowchart illustrating the operation of the sound reproduction system according to the implementation method.
[0019] Figure 4 Figure 1 illustrates the third position of the acoustic image localized using the third head-related transfer function of the implementation method.
[0020] Figure 5 This is a flowchart illustrating the operation of a modified example of an implementation method of an audio reproduction system.
[0021] Figure 6A Figure 1 illustrates the third position of the acoustic image localized by the third head-related transfer function of a modified embodiment.
[0022] Figure 6B This is Figure 2, which illustrates the third position of the acoustic image being located using the third head-related transfer function of a modified embodiment.
[0023] Figure 6C This is the third figure illustrating the third position of the acoustic image localized by the third head-related transfer function of the modified embodiment. Detailed Implementation
[0024] (Based on the understanding that forms the basis of this disclosure)
[0025] Previously, techniques were known to enable users to perceive stereo sound by controlling the position of the sound image of a sound source object perceived by the user within a virtual three-dimensional space (hereinafter referred to as a three-dimensional sound field) to achieve stereo sound reproduction (see, for example, Patent Document 1). By positioning the sound image at a predetermined location within the virtual three-dimensional space, the user can perceive the sound as if it were emanating from that predetermined location. In order to position the sound image at a predetermined location within the virtual three-dimensional space in this way, it is necessary, for example, to perform calculations on the collected sound to generate the time difference of sound arrival between the two ears and the difference in sound level between the two ears, as if the sound were perceived as stereo.
[0026] As an example of such computational processing, there is a known process of convolving the head-related transfer function (HRTF), used to perceive sound as arriving from a specified location, with the target sound signal. By performing this HTF convolution process at higher resolution, the sense of presence in the user experience is enhanced. On the other hand, HTF convolution is a computationally intensive process and requires significant computational resources. That is, to perform the HTF convolution process at high resolution, high-performance computing devices and the associated power consumption are required.
[0027] Furthermore, the development of virtual reality (VR) technology has been actively advancing in recent years. The main goal of VR is to make the position of the virtual three-dimensional space independent of the user's movement, allowing the user to experience what it's like to move within a virtual space. In particular, this VR technology attempts to enhance the sense of presence by incorporating auditory elements into visual elements. For example, when a sound image is positioned in front of the user, if the user faces right, the sound image moves to the user's left, and if the user faces left, the sound image moves to the user's right. In this way, relative to the user's movement, the positioning of the sound image within the virtual space needs to move in the opposite direction to the user's movement.
[0028] To enhance the sense of presence in virtual space, it is necessary to increase spatial resolution to implement convolution processing of the head-related transfer function. Therefore, in virtual reality and similar applications, the constraints on computing devices and power consumption become more pronounced in order to achieve sound reproduction that allows users to perceive stereo sound with a high degree of presence.
[0029] Therefore, in view of the above-mentioned problems, this disclosure implements more appropriate computational processing by reducing the computational processing load while suppressing the reduction in the sense of presence. The object of this disclosure is to provide a method for sound reproduction that enables the user to perceive stereo sound through such appropriate computational processing.
[0030] More specifically, the sound reproduction method of the present disclosure enables a user to perceive a first sound as arriving from a first position in a three-dimensional sound field, and enables the user to perceive a second sound as arriving from a second position different from the first position. The method includes: an acquisition step of acquiring the movement speed of the user's head; and a generation step of generating an output sound signal for enabling the user to perceive a sound arriving from a predetermined position in the three-dimensional sound field. In the generation step, if the acquired movement speed is greater than a first threshold, the output sound signal is generated for enabling the user to perceive both the first and second sounds as arriving from a third position between the first and second positions.
[0031] According to this sound reproduction method, when the user's head movement speed is greater than a first threshold, the first sound perceived as arriving from a first position and the second sound perceived as arriving from a second position can be perceived as arriving from a third position. In this case, the processing for localizing the sound image of the first sound to the first position and the processing for localizing the sound image of the second sound to the second position can be commonplaced into processing for localizing to the third position, thus reducing the processing load. Furthermore, if the first threshold is set to a value that blurs the user's perception of the sound image position when the user's head movement speed exceeds this threshold, the impact of changes in sound image position on the sense of presence can be suppressed even after the aforementioned processing. This also reduces the user's sense of dissonance that may result from reduced processing load. Therefore, more appropriate computational processing allows the user to perceive stereo sound.
[0032] Alternatively, in the above generation step, if the obtained motion speed is below the first threshold, the output sound signal is generated by convolving the first head-related transfer function used to locate the sound to the first position with the first sound signal related to the first sound, and by convolving the second head-related transfer function used to locate the sound to the second position with the second sound signal related to the second sound. If the obtained motion speed is greater than the first threshold, the output sound signal is generated by convolving the third head-related transfer function used to locate the sound to the third position with the summed sound signal obtained by adding the first sound signal to the second sound signal.
[0033] When positioning the sound image of the first sound at the first position, the first head-related transfer function is convolved with the first sound signal related to the first sound. When positioning the sound image of the second sound at the second position, the second head-related transfer function is convolved with the second sound signal related to the second sound. According to the above description, when positioning the sound images of the first and second sounds at the third position, only the convolution of the summed sound signal (obtained by adding the first and second sound signals) with the third head-related transfer function used to position the sound at the third position is required. That is, the convolution processing of the first head-related transfer function for the first sound signal and the convolution processing of the second head-related transfer function for the second sound signal can be commonplaced into the convolution processing of the third head-related transfer function of the summed sound signal. This reduces the processing load, allowing the user to perceive stereo sound through more appropriate computational processing.
[0034] Alternatively, for example, the aforementioned motion speed could be the rotational speed of the user's head rotating around a first axis passing through the user's head, and the aforementioned third position could be a position on a bisecting line in a virtual plane viewed from the direction of the first axis, bisecting the angle between the straight lines connecting the first position and the second position to the user.
[0035] Therefore, the set third position can be used in accordance with the user's head rotation. The third position is set on the bisecting line of the line connecting the first and second positions to the user, within a virtual plane viewed from the direction of the first axis (which serves as the rotation axis). Thus, the third position can be set to the direction between the direction of the first position and the direction of the second position as observed by the user, to match the direction of sound arrival, which becomes blurred due to the user's rotation. Therefore, while reducing processing load, it is possible to suppress the dissonance of the sound arrival direction and allow the user to perceive stereo sound.
[0036] Alternatively, for example, the aforementioned rotational speed may be obtained as the amount of rotation per unit time detected by a detector that moves integrally with the user's head, detecting the amount of rotation with at least one of three mutually orthogonal axes as the rotation axis.
[0037] Therefore, the rotation speed of the user's head can be obtained using a detector as the motion speed. Thus, based on the rotation speed obtained as described above, the dissonance of the sound's direction of arrival can be suppressed, allowing the user to perceive stereo sound.
[0038] Alternatively, for example, the aforementioned movement speed may be the displacement speed of the user's head along the direction of the second axis passing through the user's head, and the displacement speed may be obtained as the amount of displacement per unit time detected by the detector, which moves integrally with the user's head and detects the amount of displacement with at least one of the three mutually orthogonal axes as the displacement direction.
[0039] The set third position can be used in accordance with the movement of the user's head. At this time, the displacement velocity of the user's head can be obtained using a detector. Therefore, based on the displacement velocity obtained as described above, the dissonance of the sound's direction of arrival can be suppressed, allowing the user to perceive stereo sound.
[0040] Alternatively, in the above-described sound reproduction method, the user may perceive multiple sounds, which are sounds arriving from various locations within a defined area on the three-dimensional sound field, including the first and second locations, and at least the first and second sounds. In the generation step, if the movement speed is greater than the first threshold, an output sound signal is generated to make the user perceive all of the multiple sounds as sounds arriving from the third location.
[0041] Therefore, users can perceive all multiple sounds within a specified range as originating from the third position. Thus, the head correlation transfer function used to localize the sound image to the third position can be commonalized to convolve the sounds within the specified range individually. Consequently, the processing complexity of convolving the head correlation transfer function is reduced, allowing users to perceive stereo sound through more appropriate computational processing.
[0042] Alternatively, for example, in the above-described sound reproduction method, the user may perceive the first intermediate sound as a sound arriving from a first intermediate position between the first position and the third position, and the user may perceive the second intermediate sound as a sound arriving from a second intermediate position between the second position and the third position. In the above-described generation step, if the movement speed is below the first threshold and above a second threshold that is less than the first threshold, the above-described output sound signal is generated to make the user perceive the first intermediate sound and the second intermediate sound as a sound arriving from the third position.
[0043] Therefore, the same processing described above can be applied within a narrower range, including the first and second intermediate positions, which are closer to the third position than the first and second positions, respectively. Here, since the user's head movement speed is less than the first threshold, if the sound from the first and second positions is concentrated in the third position, a change in the sound image position will be perceived, potentially leading to a sense of dissonance; therefore, this processing is not implemented. On the other hand, since the user's head movement speed is greater than the second threshold, even if the sound from a narrow range, narrower than the specified range including the first and second positions, is concentrated in the third position, a change in the sound image position will not be perceived. Therefore, when the movement speed is below the first threshold and greater than the second threshold (which is less than the first threshold), the sound from the first and second intermediate positions within such a narrow range can be concentrated in the third position, reducing the computational processing load. Thus, more appropriate computational processing allows the user to perceive stereo sound.
[0044] Furthermore, the sound reproduction system of one of the technical solutions disclosed herein enables a user to perceive a first sound as a sound arriving from a first position in a three-dimensional sound field, and enables the user to perceive a second sound as a sound arriving from a second position different from the first position. This system includes: an acquisition unit that acquires the movement speed of the user's head; and a generation unit that generates an output sound signal for enabling the user to perceive a sound arriving from a predetermined position in the three-dimensional sound field. If the acquired movement speed is greater than a first threshold, the generation unit generates the output sound signal for enabling the user to perceive both the first and second sounds as sounds arriving from a third position between the first and second positions.
[0045] Thus, an audio reproduction system can be achieved that has the same effect as the audio reproduction method described above.
[0046] Furthermore, the technical solution disclosed herein can also be implemented as a non-temporary recording medium that can be read by a computer and contains a program for enabling a computer to execute the aforementioned sound reproduction method.
[0047] Therefore, a computer can achieve the same effect as the sound reproduction method described above.
[0048] Furthermore, these inclusive or specific technical solutions can also be implemented by non-transitory recording media such as systems, devices, methods, integrated circuits, computer programs, or computer-readable CD-ROMs, or by any combination of systems, devices, methods, integrated circuits, computer programs, and recording media.
[0049] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Furthermore, the embodiments described below are inclusive or specific examples. The numerical values, shapes, materials, constituent elements, arrangement and connection methods of constituent elements, steps, and order of steps shown in the following embodiments are examples and are not intended to limit this disclosure. In addition, constituent elements in the following embodiments that are not described in the independent claims are described as arbitrary constituent elements. Furthermore, the figures are schematic diagrams and are not necessarily strictly illustrated. Also, in the figures, substantially identical structures are given the same reference numerals, and repeated descriptions may be omitted or simplified.
[0050] Furthermore, in the following description, elements are sometimes assigned ordinal numbers such as 1st, 2nd, and 3rd. These ordinal numbers are assigned to elements for identification purposes and do not necessarily correspond to a meaningful order. These ordinal numbers may also be appropriately replaced, reassigned, or removed.
[0051] (Implementation Method)
[0052] [summary]
[0053] First, an overview of the sound reproduction system of the implementation method will be given. Figure 1 This is a schematic diagram illustrating an example of the use of the sound reproduction system according to the implementation method. Figure 1 The data indicates that 99 users are using the audio playback system 100.
[0054] Figure 1 The sound reproduction system 100 and the stereoscopic image reproduction system 200 shown are used simultaneously. As explained above, in this embodiment, by simultaneously viewing and listening to stereoscopic images and stereoscopic sounds, the auditory presence of the image is enhanced, and the visual presence of the sound is enhanced, allowing the user to experience as if they were at the scene where the image and sound were being captured. For example, it is known that when displaying an image (moving image) of people conversing, even if the location of the sound image of the conversation deviates from the corner of the person's mouth, the user 99 will still perceive the conversation as coming from that person's mouth. By correcting the position of the sound image based on visual information, the presence of the image and sound can be enhanced together.
[0055] The stereoscopic image reproduction system 200 is an image display device worn on the head of the user 99. Therefore, the stereoscopic image reproduction system 200 moves as a unit with the user 99's head. For example, as shown in the illustration, the stereoscopic image reproduction system 200 is a glasses-type device supported by the user 99's ears and nose.
[0056] The stereoscopic image reproduction system 200 changes the displayed image according to the movement of the user 99's head, making the user 99 perceive it as if their head is moving within a three-dimensional image space. That is, when an object in the three-dimensional image space is in front of the user 99, if the user 99 is facing right, the object moves to the left of the user 99; if the user 99 is facing left, the object moves to the right of the user. In this way, relative to the user 99's movement, the stereoscopic image reproduction system 200 moves the three-dimensional image space in the opposite direction to the user 99's movement.
[0057] The stereoscopic image reproduction system 200 displays two images with a disparity in the amount of parallax to the left and right eyes of the user 99, respectively. Based on the disparity in the displayed images, the user 99 can perceive the three-dimensional position of objects in the images. Furthermore, when the user 99 uses the sound reproduction system 100 with their eyes closed, for example, to reproduce therapeutic sounds for sleep induction, it is not necessary to use the stereoscopic image reproduction system 200 simultaneously. That is, the stereoscopic image reproduction system 200 is not an essential component of this disclosure.
[0058] The sound reproduction system 100 is an audio prompting device worn on the head of the user 99. Therefore, the sound reproduction system 100 moves as a unit with the user 99's head. For example, the sound reproduction system 100 consists of two earbud-type devices worn independently on the left and right ears of the user 99. These two devices communicate with each other to synchronously provide prompts for the right and left ears.
[0059] The sound reproduction system 100 changes the audible prompts according to the movement of the user 99's head, making the user 99 perceive it as if the user 99 is moving their head within a three-dimensional sound field. Therefore, as described above, relative to the user 99's movement, the sound reproduction system 100 causes the three-dimensional sound field to move in the opposite direction to the user's movement.
[0060] Here, it is known that if the user 99's head movement becomes constant, the user 99's recognition of the location of the sound image within the three-dimensional sound field becomes blurred. The sound reproduction system 100 of this embodiment utilizes this phenomenon to reduce the computational processing load. That is, the sound reproduction system 100 obtains the speed of the user 99's head movement, and if the obtained speed is greater than a first threshold, it makes multiple sounds perceived as arriving from a predetermined area in the three-dimensional sound field perceive as arriving from a single point within that predetermined area.
[0061] This defined area corresponds to the range where the user's perception of the sound image location becomes blurred due to the rapid speed of head movement. Therefore, this defined area needs to be set for each user, which can be done, for example, through prior experimentation. Furthermore, since the defined area is also affected by the amount of head movement of the user, it can also be set by detecting the amount of head movement of the user to determine the defined area corresponding to the amount of movement.
[0062] Furthermore, regarding the first threshold for movement speed, it is also necessary to define a user-specific value at which the user's perception of sound image location becomes blurred due to a certain level of movement speed. Therefore, a value set through prior experiments can be used. Alternatively, a generalized defined area and the first threshold can be set by averaging the experimental results from multiple users.
[0063] [structure]
[0064] Next, refer to Figure 2 The structure of the sound reproduction system 100 of this embodiment will be described. Figure 2 This is a block diagram illustrating the functional structure of the sound reproduction system according to the implementation method.
[0065] like Figure 2 As shown, the sound reproduction system 100 of this embodiment includes a processing module 101, a communication module 102, a detector 103, and a driver 104.
[0066] The processing module 101 is a computing device used to perform various signal processing in the sound reproduction system 100. The processing module 101 includes, for example, a processor and a memory, and performs various functions by having the processor execute programs stored in the memory.
[0067] The processing module 101 includes an input unit 111, an acquisition unit 121, a generation unit 131, and an output unit 141. Details of each functional unit in the processing module 101 will be described below along with details of other structures of the processing module 101.
[0068] The communication module 102 is an interface device used to receive audio signals input to the audio playback system 100. The communication module 102 includes, for example, an antenna and a signal converter, and receives audio signals from external devices via wireless communication. More specifically, the communication module 102 uses the antenna to receive wireless signals representing audio signals that have been converted into a form suitable for wireless communication, and then performs a further conversion from wireless signals to audio signals using the signal converter. Thus, the audio playback system 100 obtains audio signals from external devices via wireless communication. The audio signals obtained by the communication module 102 are input to the input unit 111. In this way, the audio signals are input to the processing module 101. Alternatively, communication between the audio playback system 100 and external devices can also be performed via wired communication.
[0069] The audio signal acquired by the audio reproduction system 100 is encoded in a specified format, such as MPEG-H Audio. For example, the encoded audio signal contains information related to the sound to be reproduced by the audio reproduction system 100, and information related to the positioning position of the sound image within a specified location in a three-dimensional sound field. For instance, the audio signal contains information related to multiple sounds, including a first sound and a second sound, so that the sound images of each sound are positioned at different locations within the three-dimensional sound field when reproduced.
[0070] This stereo sound, for example, together with an image recognized by the stereoscopic image reproduction system 200, can enhance the sense of presence of audiovisual content. Furthermore, the sound signal may contain only information about the sound. In this case, information about the location can also be obtained separately. Moreover, as described above, the sound signal may contain a first sound signal about a first sound and a second sound signal about a second sound, but it is also possible to obtain multiple sound signals, each containing only the first or second sound signal, and reproduce them simultaneously to position the sound image at different locations within the three-dimensional sound field. In this way, the form of the input sound signal is not particularly limited, as long as the sound reproduction system 100 has an input unit 111 corresponding to various forms of sound signals.
[0071] Detector 103 is a device used to detect the movement speed of the user 99's head. Detector 103 is constructed by combining various sensors used for motion detection, such as a gyroscope sensor and an accelerometer. In this embodiment, detector 103 is built into the audio playback system 100, but it can also be built into an external device, such as a stereoscopic image playback system 200 that operates according to the movement of the user 99's head, just like the audio playback system 100. In this case, detector 103 may not be included in the audio playback system 100. Furthermore, as detector 103, an external camera device or the like can be used to capture the movement of the user 99's head, and the movement of the user 99 can be detected by processing the captured images.
[0072] The detector 103 is integrally fixed to the housing of the sound reproduction system 100, for example, and detects the movement speed of the housing. The sound reproduction system 100 moves integrally with the user's head after being worn by the user 99, so it is possible to detect the movement speed of the user's head.
[0073] Detector 103 can, for example, detect the amount of rotation in three-dimensional space with at least one of three mutually orthogonal axes as the rotation axis as the amount of motion of the user 99's head, or it can detect the amount of displacement with at least one of the three axes as the displacement direction as the amount of motion of the user 99's head. Furthermore, detector 103 can also detect both the amount of rotation and the amount of displacement as the amount of motion of the user 99's head.
[0074] The acquisition unit 121 acquires the head movement speed of the user 99 from the detector 103. More specifically, the acquisition unit 121 acquires the amount of head movement of the user 99 detected by the detector 103 per unit time as the movement speed. In this way, the acquisition unit 121 acquires at least one of the rotational speed and the displacement speed from the detector 103.
[0075] Here, the generation unit 131 determines whether the obtained head movement speed of the user 99 is greater than the first threshold mentioned above. Based on the result of this determination, the generation unit 131 decides whether to reduce the computational processing load. More detailed operation of the generation unit 131 will be described later. According to the above decision, the generation unit 131 performs computational processing on the input sound signal to generate an output sound signal for use as a prompt.
[0076] Output unit 141 is a functional unit that outputs the generated sound signal to driver 104. Driver 104 performs signal conversion from digital to analog based on the output sound signal, thereby generating a waveform signal. Sound waves are generated based on the waveform signal to provide a sound prompt to user 99. Driver 104 includes, for example, a diaphragm, a magnet, and a voice coil drive mechanism. Driver 104 actuates the drive mechanism according to the waveform signal, causing the diaphragm to vibrate. Thus, driver 104 generates sound waves through the vibration of the diaphragm corresponding to the output sound signal. The sound waves propagate through the air and reach the ear of user 99, allowing user 99 to perceive the sound.
[0077] [action]
[0078] Next, refer to Figure 3 The operation of the sound reproduction system 100 described above will be explained. Figure 3 This is a flowchart illustrating the operation of the sound reproduction system according to the implementation method. For example... Figure 3As shown, firstly, if the operation of the sound reproduction system 100 is started, a first sound signal relating to the first sound and a second sound signal relating to the second sound are acquired (step S101). Here, the processing module 101 acquires the sound signal containing the first sound signal and the second sound signal by inputting the sound signal obtained from an external device by the communication module 102 to the input unit 111.
[0079] Next, the acquisition unit 121 acquires the head movement speed of the user 99 from the detector 103 as a detection result (acquisition step S102). The generation unit 131 compares the acquired movement speed with a first threshold and determines whether the movement speed is greater than the first threshold (step S103). If the movement speed is below the first threshold ("No" in step S103), the sound reproduction system 100 causes the user 99 to perceive the first sound and the second sound as sounds arriving from their respective original sound image positions, namely the first position and the second position. Therefore, the generation unit 131 convolves the first sound signal with a first head correlation transfer function used to locate the sound image to the first position. In addition, the generation unit 131 convolves the second sound signal with a second head correlation transfer function used to locate the sound image to the second position (step S104). The generation unit 131 generates an output sound signal containing the first sound signal and the second sound signal that have undergone such convolution processing (step S105).
[0080] On the other hand, when the movement speed is greater than the first threshold ("Yes" in step S103), the sound reproduction system 100 causes the user 99 to perceive the sound as arriving from a third position between the original sound image positions of the first and second sounds, i.e., the first and second positions. Therefore, the generation unit 131 generates an additive sound signal related to the sound after the first and second sounds are superimposed by adding the first sound signal and the second sound signal. In addition, the area between the first and second positions refers, for example, to the region enclosed by a virtual straight line passing through the first position and other virtual straight lines parallel to that virtual straight line and passing through the second position. At this time, the area may also include the virtual straight line and other virtual straight lines.
[0081] The generation unit 131 then convolves the summed audio signal with the third head correlation transfer function used to locate the sound image at the third position (step S107). The generation unit 131 generates an output audio signal containing the summed audio signal that has undergone such convolution processing (step S108). Alternatively, steps S103 to S108 can be collectively referred to as the generation steps.
[0082] The output unit 141 drives the driver 104 by outputting the output sound signal generated by the generation unit 131 to the driver 104, thereby prompting a sound based on the output sound signal (step S106). In this way, the first sound and the second sound can be perceived together as sounds arriving from the third position, thus simplifying the computational processing used for sound image localization compared to the case where the first sound is perceived as arriving from the first position and the second sound is perceived as arriving from the second position. This temporarily reduces the request processing capacity, and reduces heat generated by the processor and power consumption associated with computational processing. Furthermore, as described above, the simplification of computational processing also blurs the user's perception of sound image location, thus minimizing the impact on the sense of presence. In the sound reproduction system 100, the computational processing can be simplified as needed, allowing the user to perceive stereo sound through more appropriate computational processing.
[0083] Here, refer to Figure 4 The third point above will be explained in more detail. Figure 4 This is a diagram illustrating the third position of acoustic image localization using the third head-related transfer function of the implementation method. Additionally, in Figure 4 In the diagram, black dots represent the positions of sound images within the three-dimensional sound field, and arrows extending from the black dots towards user 99 indicate the direction from which the sound is coming towards user 99. Additionally, virtual speakers are also represented at the black dots indicating the positions of the sound images.
[0084] exist Figure 4 In the example shown, it is assumed that user 99 is rotating their head at a speed greater than the first threshold. Alternatively, the following action can be performed if user 99 displaces their head at a speed greater than the first threshold. In this example, as indicated by the hollow double-headed arrow, user 99's head rotates around a first axis perpendicular to the plane of the paper. In this case, as shown in the figure, the third position P3 or P3a is located on the bisecting line of the angle formed by the line connecting the first position P1 or P1a to user 99 and the line connecting the second position P2 or P2a to user 99. The bisecting line is indicated by a dotted-shaded arrow in the figure.
[0085] In this way, by simplifying the convolution calculation of the head-related transfer function, more appropriate calculation processing can enable the user to perceive stereo sound. Furthermore, if the head-related transfer function contains information related to the distance at which the sound image is located, it can be configured such that multiple head-related transfer functions are prepared to locate the sound image at multiple distances in the same direction of sound arrival, and one of these head-related transfer functions is convolved. In this case, since the directions of arrival of the first and second sounds and their distances to the sound image positions are averaged, the user is more likely to perceive dissonance. Therefore, structures such as setting a smaller defined area can also be included to reduce the perception of dissonance.
[0086] In the case of user 99 displacing their head, let's assume the displacement velocity is greater than a first threshold. In this example, user 99's head might displace along a second axis in the vertical direction along the plane of the paper. In this case, the third position P3 is a position on an equidistant line orthogonal to the second axis and equidistant from both the first position P1 and the second position P2. By positioning the acoustic image at such a position, an average third position P3 can be set in the region where distances become blurred due to the user 99's head displacement. Alternatively, the direction of user 99's head displacement could also be a single direction.
[0087] Furthermore, when setting the third position, it can also be a position corresponding to either the first or second position. For example, if the first sound is a person's dialogue in the content, and the second sound is ambient sound in the content, the first sound takes priority, and the sound image position set for the first sound is set as the third position. Thus, the first and second sounds are perceived as sounds arriving from the first position, which is set as the third position. In this case, the first head-related transfer function, which is used to make the user perceive the sound as arriving from the first position, is directly used.
[0088] That is, in this example, the head-related transfer function that has already been used is used, so it is not necessary, for example, to set a position that does not correspond to either the first or second position originally set by the sound signal as the third position, as shown in the example above. In other words, the sound image position originally set by the sound signal can be set as the third position. Therefore, the head-related transfer function used to locate the sound image to the originally set sound image position can be used, so it is not necessary to use mapping information, which maps the head-related transfer function used to make the user 99 perceive the sound as arriving from any point in the three-dimensional sound field. Therefore, the determination process of the head-related transfer function for the set third position is simplified, and the user 99 can perceive stereo sound through more appropriate calculation processing. In this way, the range between the first and second positions includes the first and second positions themselves.
[0089] Furthermore, the third position can be set either at the midpoint of the line segment connecting the first and second positions in space, or simply at a random position between the first and second positions.
[0090] [Variation Example]
[0091] The following is for reference Figure 5 Figure 6 illustrates the operation of the sound reproduction system of a modified embodiment. Furthermore, in the following description of the modified embodiment, the differences from the above-described embodiment will be the focus of the explanation, while substantially equivalent points will be omitted or simplified.
[0092] Figure 5 This is a flowchart illustrating the operation of a modified example of an implementation method of an audio reproduction system. Figure 6A Figure 1 illustrates the third position of the acoustic image localized by the third head-related transfer function of a modified embodiment. Figure 6B This is Figure 2, which illustrates the third position of the acoustic image being located using the third head-related transfer function of a modified embodiment. Figure 6C This is the third diagram illustrating the third position of the sound image localized by the third head-correlation transfer function in a modified embodiment. The sound reproduction system of this modified embodiment differs from the sound reproduction system 100 of the above-described embodiment in that it convolves the head-correlation transfer function with the sound signal, using the first and second thresholds as boundaries, to represent the change in the target sound.
[0093] More specifically, in the sound reproduction system of this variant, a second threshold smaller than the first threshold is set. The first threshold, similar to the implementation described above, is used to determine whether to apply a third head correlation transfer function to make the user 99 perceive the first and second sounds as sounds arriving from the third position. In this variant, by using the determination of the second threshold, convolution is performed to make the user 99 perceive the first intermediate sound and the second intermediate sound, located at a first intermediate position and a second intermediate position closer to the third position than the first and second sounds, as sounds arriving from the third position using the third head correlation transfer function, thereby reducing the amount of computational processing.
[0094] Here, the movement speed of the user's head (99) is determined. If the movement speed is below the second threshold, the first sound is located at position 1 P1, the second sound is located at position 2 P2, and the first intermediate sound is located at position 1 P1m (refer to...). Figure 6A (etc.), the second middle sound is located at the second middle position P2m (refer to) Figure 6A(etc.). On the other hand, when the head movement speed of user 99 is greater than the first threshold, as described above, the third head correlation transfer function is applied to the audio signals of the first and second sounds (i.e., the first and second audio signals). At this time, the third head correlation transfer function is also applied to the audio signals of the first and second intermediate sounds (i.e., the first and second intermediate audio signals), and the first sound, the second sound, the first intermediate sound, and the second intermediate sound are all located at the third position P3.
[0095] In addition, in this variation, when the head movement speed of user 99 is greater than the second threshold but less than the first threshold, the first sound is located at position 1 P1, the second sound is located at position 2 P2, and the first intermediate sound and the second intermediate sound are located at position 3 P3. That is, in this variation, when the head movement speed of user 99 is not so fast, such as below the second threshold, the calculation of the convolution of the head-related transfer function is simplified for a smaller defined region (i.e., a narrow region) that does not include positions 1 P1 and 2 P2 but includes positions 1 intermediate P1m and 2 intermediate P2m.
[0096] As an example of the operation of the sound reproduction system in this variation, such as Figure 5 As shown, after the acquisition unit 121 acquires the motion speed (step S102), the generation unit 131 determines whether the motion speed is greater than a second threshold (step S201). If the motion speed is below the second threshold ("No" in step S201), the process proceeds to step S202, where, similar to the embodiment described above, the operation of convolving each sound signal with a head correlation transfer function to locate the sound image to the position it should be located is performed (step S202). That is, for the first sound signal relating to the first sound, a first head correlation transfer function is convolved to locate the sound image to the first position P1; for the second sound signal relating to the second sound, a second head correlation transfer function is convolved to locate the sound image to the second position P2; for the first intermediate sound signal relating to the first intermediate sound, a first intermediate head correlation transfer function is convolved to locate the sound image to the first intermediate position P1m; and for the second intermediate sound signal relating to the second intermediate sound, a second intermediate head correlation transfer function is convolved to locate the sound image to the second intermediate position P2m.
[0097] On the other hand, if the movement speed is greater than the second threshold (yes in step S201), the generation unit 131 also determines whether the movement speed is greater than the first threshold (step S204). If the movement speed is less than the first threshold (no in step S204), the sound reproduction system 100 causes the user 99 to perceive the first intermediate sound and the second intermediate sound as sounds arriving from the third position. Therefore, the generation unit 131 convolves the third head correlation transfer function into the summed sound signal obtained by adding the first intermediate sound signal related to the first intermediate sound and the second intermediate sound signal related to the second intermediate sound (step S205). The generation unit 131 generates an output sound signal that includes the first sound signal, the second sound signal, and the summed sound signal obtained by adding the first intermediate sound signal and the second intermediate sound signal (step S206). Then, proceed to step S106 and perform the same operation as in the above-described embodiment.
[0098] On the other hand, if the movement speed is greater than the first threshold ("Yes" in step S204), the process proceeds to step S207, where the third head correlation transfer function is convolved into the summed sound signal obtained by adding the first and second sound signals, using the same operation as in the above embodiment. In this modified example, a first intermediate sound signal and a second intermediate sound signal are also added to the summed sound signal, and the first sound, the second sound, the first intermediate sound, and the second intermediate sound are perceived by the user 99 as sounds arriving from the third position P3.
[0099] As a result of the above actions, in the modified sound reproduction system of this embodiment, when the user's movement speed is below the second threshold, a sound field is formed within the three-dimensional sound field. Figure 6A The audio-visual image shown. Additionally, in Figure 6A In, with Figure 4 This also shows a diagram of the three-dimensional sound field viewed from the first axis. For example... Figure 6A As shown, when the user's movement speed is below the second threshold, the first sound, the second sound, the first intermediate sound, and the second intermediate sound are perceived by the user as sounds arriving from their original sound image positions.
[0100] Furthermore, in the sound reproduction system of this modified example, when the user's movement speed is below the first threshold and above the second threshold, a sound field is formed within the three-dimensional sound field. Figure 6B The audio-visual image shown. Additionally, in Figure 6B In, with Figure 4 The diagram also shows the three-dimensional sound field viewed from the first axis direction.
[0101] like Figure 6BAs shown, when user 99's movement speed is below the first threshold but above the second threshold, the first intermediate sound, which user 99 originally perceived as arriving from the first intermediate position P1m (closer to the third position P3 than the first position P1), is perceived by user 99 as arriving from the third position P3. Similarly, when the movement speed is below the first threshold but above the second threshold, the second intermediate sound, which user 99 originally perceived as arriving from the second intermediate position P2m (closer to the third position P3 than the second position P2), is perceived by user 99 as arriving from the third position P3.
[0102] Furthermore, in the sound reproduction system of this modified example, when the user's movement speed is greater than the first threshold, a sound field is formed within the three-dimensional sound field. Figure 6C The audio-visual image shown. Additionally, in Figure 6C In, with Figure 4 The diagram also shows the three-dimensional sound field viewed from the first axis direction.
[0103] like Figure 6C As shown, when the user's movement speed is greater than the first threshold, all the sounds that were originally located in the specified area including the first intermediate position P1m and the second intermediate position P2m and including the first position P1 and the second position P2 are perceived by the user as sounds arriving from the third position P3.
[0104] Thus, when the movement speed exceeds the second threshold, the sound within a defined area corresponding to the user's movement speed is perceived by the user as arriving from the third position P3. For example, in the diagram, when the movement speed exceeds the first threshold, the sound within the defined area represented by the long dashed line is perceived by the user as arriving from the third position P3. Furthermore, when the movement speed exceeds the second threshold but is below the first threshold, the sound within the narrow defined area (i.e., the narrow area) represented by the dashed line is perceived by the user as arriving from the third position P3.
[0105] Furthermore, the third position P3 can be considered as the first intermediate position P1m and the second intermediate position P2m. That is, the third position P3 is set based on the four positions: the first position P1, the second position P2, the first intermediate position P1m, and the second intermediate position P2m. Here, for example, the third position P3 can be set on the straight line connecting the centers of the first position P1, the second position P2, the first intermediate position P1m, and the second intermediate position P2m to the user 99, and at a distance equal to the shortest distance among the distances from each of the first position P1, the second position P2, the first intermediate position P1m, and the second intermediate position P2m to the user 99. Alternatively, the third position P3 can also be set as the average coordinate of the coordinates corresponding to the four positions in the planar coordinate system viewed from the first axis direction.
[0106] Alternatively, it can be configured to have three or more stages, such as a third threshold for the user's (99's) movement speed, so that sound within a smaller, defined area is perceived by the user (99) as arriving from the third position P3. The number of stages in the relationship between movement speed and the size of the defined area is not particularly limited.
[0107] Furthermore, regarding the second threshold, it can be set based on a value inherent to the user 99, such as the level of motion speed at which the user 99's perception of the sound image location becomes blurred, just like the first threshold described in the above-described implementation, or it can be set to a generalized value.
[0108] (Other implementation methods)
[0109] The implementation methods have been described above, but this disclosure is not limited to the implementation methods described above.
[0110] For example, in the above embodiments, an example was described where the sound does not follow the movement of the user's head, but the present disclosure is also effective in cases where the sound follows the movement of the user's head. That is, in an action where the user perceives the first sound as a sound arriving from a first position that has moved relatively with the movement of the user's head, and the user perceives the second sound as a sound arriving from a second position that has moved relatively with the movement of the user's head, if the speed of the head movement is greater than a first threshold, the first and second sounds are perceived as sounds arriving from a third position that has moved relatively with the movement of the user's head.
[0111] In this case, the head-related transfer function used to locate the first and second sounds to the first and second positions is also convolved with each sound signal. Since the head-related transfer function to be convolved with the sound signals is common to the first threshold, the computational processing is simplified. That is, similar to the above implementation, the request processing capacity can be temporarily reduced, and the heat generated by the processor and the power consumption associated with computational processing can be reduced. On the other hand, even with such simplification of computational processing, if the user's head movement speed is high, it is difficult to accurately perceive the position of the sound image, so the user's sense of incongruity in the position of the sound image is less likely to increase. Therefore, more appropriate computational processing can enable the user to perceive stereo sound.
[0112] Furthermore, the sound reproduction system described in the above embodiments can be implemented as a single device with all its components, or the functions can be distributed among multiple devices, which then work together to achieve the desired effect. In the latter case, an information processing device such as a smartphone, tablet computer, or PC can be used as the device corresponding to the processing module.
[0113] Furthermore, the audio reproduction system disclosed herein can also be implemented as an audio processing device that is connected to a reproduction device that only has a driver, and outputs an output audio signal to the reproduction device based on a convolutional processing of the acquired audio signal using a head correlation transfer function. In this case, the audio processing device can be implemented either as hardware with dedicated circuitry or as software to enable a general-purpose processor to perform specific processing.
[0114] Furthermore, in the above embodiments, the processing to be performed by a specific processing unit may be performed by other processing units. Additionally, the order of multiple processes may be changed, or multiple processes may be executed in parallel.
[0115] Furthermore, in the above embodiments, each component can also be implemented by executing a software program suitable for each component. Each component can also be implemented by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.
[0116] Furthermore, each component can also be implemented in hardware. For example, each component can be a circuit (or integrated circuit). These circuits can either form a single circuit as a whole, or they can be separate circuits. Moreover, these circuits can be either general-purpose circuits or special-purpose circuits.
[0117] In addition, these inclusive or specific technical solutions can also be implemented by systems, devices, methods, integrated circuits, computer programs or computer-readable CD-ROMs and other recording media, or by any combination of systems, devices, methods, integrated circuits, computer programs and recording media.
[0118] For example, this disclosure can also be implemented as a method for reproducing sound signals executed by a computer, or as a program for causing a computer to perform the method for reproducing sound signals. This disclosure can also be implemented as a computer-readable, non-transitory recording medium containing such a program.
[0119] In addition, this disclosure also includes forms obtained by applying various modifications to each embodiment that can be conceived by those skilled in the art, or forms achieved by arbitrarily combining the constituent elements and functions of each embodiment without departing from the spirit of this disclosure.
[0120] Industrial availability
[0121] This disclosure is useful in the reproduction of stereo sound that is perceived by the user in relation to the movement of the user's head.
[0122] Label Explanation
[0123] 99 users
[0124] 100-sound reproduction system
[0125] 101 Processing Module
[0126] 102 Communication Module
[0127] 103 detector
[0128] 104 drives
[0129] 111 Input Section
[0130] 121 Acquisition Department
[0131] 131 Generation Department
[0132] 141 Output Section
[0133] 200 Stereoscopic Image Reproduction System
[0134] P1, P1a, Position 1
[0135] P2, P2a, second position
[0136] P3, P3a, Position 3
[0137] P1m, first middle position
[0138] P2m is the second middle position.
Claims
1. A method for sound reproduction, wherein a user perceives a first sound as arriving from a first position in a three-dimensional sound field, and the user perceives a second sound as arriving from a second position different from the first position, wherein, include: The step is to obtain the head movement speed of the aforementioned user; as well as The generation step generates an output sound signal used to enable the user to perceive sound arriving from a specified position in the three-dimensional sound field. In the above generation step, if the obtained motion speed is greater than the first threshold, the above output sound signal is generated to make the user perceive the first sound and the second sound as sounds arriving from a third position between the first position and the second position.
2. The sound reproduction method as described in claim 1, wherein, In the above generation steps, If the obtained motion speed is below the first threshold, the output sound signal is generated by convolving the first head-related transfer function used to locate the sound to the first position with a first sound signal relating to the first sound, and by convolving the second head-related transfer function used to locate the sound to the second position with a second sound signal relating to the second sound. If the obtained motion speed is greater than the first threshold, the output sound signal is generated by convolving the third head-related transfer function used to locate the sound to the third position with the summed sound signal obtained by adding the first sound signal to the second sound signal.
3. The sound reproduction method as described in claim 1, wherein, The aforementioned speed of motion is the rotational speed of the user's head as it rotates around the first axis passing through the user's head. The aforementioned third position is located on the bisecting line of the line that bisects the angle between the lines connecting the aforementioned first position and the aforementioned second position to the aforementioned user, within a virtual plane viewed from the direction of the aforementioned first axis.
4. The sound reproduction method as described in claim 3, wherein, The aforementioned rotational speed is obtained as the amount of rotation per unit time detected by the detector, which moves integrally with the user's head and detects the amount of rotation with at least one of three mutually orthogonal axes as the rotation axis.
5. The sound reproduction method as described in claim 1, wherein, The aforementioned velocity is the displacement velocity of the user's head along the second axis passing through the user's head. The aforementioned displacement velocity is obtained as the amount of displacement per unit time detected by the detector, which moves integrally with the user's head and detects the amount of displacement with at least one of three mutually orthogonal axes as the displacement direction.
6. The sound reproduction method as described in claim 1, wherein, In the above-described sound reproduction method, the user perceives multiple sounds, which are sounds arriving from various locations within a defined area on the three-dimensional sound field, including the first and second positions, and include at least the first and second sounds. In the above generation step, when the above movement speed is greater than the above first threshold, the above output sound signal is generated to make the above user perceive all of the above multiple sounds as sounds arriving from the above third position.
7. The sound reproduction method according to any one of claims 1 to 6, wherein, In the above-described sound reproduction method, the user perceives the first intermediate sound as a sound arriving from a first intermediate position between the first position and the third position, and the user perceives the second intermediate sound as a sound originating from a second intermediate position between the second position and the third position. In the above generation step, if the movement speed is below the first threshold and greater than the second threshold which is less than the first threshold, the above output sound signal is generated to make the user perceive the first intermediate sound and the second intermediate sound as sounds arriving from the third position.
8. A recording medium, wherein, The system contains a program for causing a computer to perform the sound reproduction method of claim 1.
9. A sound reproduction system that causes a user to perceive a first sound as arriving from a first position in a three-dimensional sound field, and causes the user to perceive a second sound as arriving from a second position different from the first position, wherein, include: The acquisition unit obtains the head movement speed of the aforementioned user; as well as The generation unit generates an output sound signal to enable the user to perceive sound arriving from a predetermined position in the three-dimensional sound field. When the obtained motion speed is greater than the first threshold, the generation unit generates the output sound signal to make the user perceive the first sound and the second sound as sounds arriving from a third position between the first position and the second position.
Citation Information
Patent Citations
Voice generation program in virtual space, generation method of quadtree, and voice generation device
JP2020018620A
Binaural headphone rendering with head tracking
CN107018460A
Sound processing device and control method of sound processing device
JP2014158151A