Audio processing device and program thereof

The audio processing device and program address the issue of important audio information being obscured in 6DoF content by manipulating the spatial perception of sound objects, ensuring reliable transmission of critical audio information to the listener.

JP2025072927APending Publication Date: 2025-05-12NIPPON HOSO KYOKAI
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023183416
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-10-25
Publication Date
2025-05-12

Smart Images

  • Figure 2025072927000001_ABST
    Figure 2025072927000001_ABST
Patent Text Reader

Abstract

To provide an audio processing device and a program thereof that can reliably transmit audio information outside of a content to a listener using an object-based audio system.SOLUTION: A target direction setting unit sets a first type target direction, which is a target direction based on a listening position for a first type audio object included in the content, and a plurality of second type target directions based on the listening position for a second type audio object, which is an audio object not included in the content, sets a sum of the directional vectors of the second type target directions to zero, and updates the second type target direction such that the minimum value becomes larger when the minimum value of the angle between the first type target direction and the second type target direction is smaller than a predetermined angle reference value.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present application relates to an audio processing device and a program thereof, for example, a technology for reliably playing audio information of high importance or urgency without being hindered by audio objects of 6DoF (Degrees of Freedom) content that allows the listener to set their listening position. [Background technology]

[0002] In recent years, audio technology that supports 6DoF (see Non-Patent Documents 6 and 7) has been developed by extending object-based audio systems (see Non-Patent Documents 3-5) that can manipulate audio content that is composed of audio signals and audio metadata (see Non-Patent Documents 1 and 2). 6DoF refers to the degree of freedom of an object's movement in six directions in a three-dimensional space. In 6DoF content, the listener sets an arbitrary position and orientation, and the content viewed at the set position and orientation is simulated. It differs from conventional 3D (Dimensional) audio in that the reproduced sound changes depending on the set position or orientation. [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] ITU-R BS.2076-1, Audio Definition Model, June 2017 [Non-Patent Document 2] ITU-R BS.2125-0, A serial representation of the Audio Definition Model, January 2019 [Non-Patent Document 3] ISO / IEC 23008-3:2019, Information technology - High efficiency coding and media delivery in heterogeneous environments - Part 3 3D audio, 2019 [Non-Patent Document 4] ETSI TS 103 190-2, Digital Audio Compression (AC-4) Standard; Part2: immersive and personalized audio, V1.2.1, 2018-02 [Non-Patent Document 5] ATSC Standard: A / 342:2021 Part 3, MPEG-H System, 11 March 2021 [Non-Patent Document 6] MPEG-I Immersive Audio Encoder Input Format, Version 5, April 4, 2023 [Non-Patent Document 7] Report ITU-R BT.2420-5 (09 / 2022), Collection of usage scenarios of advanced immersive sensory media systems (09 / 2022) [Non-Patent Document 8] Katsumi Nakabayashi, Fundamental Experiments on the Interaction between Stereo Sound Images and Television Sound Images, Proceedings of the 2010 Conference of the Acoustical Society of Japan, Vol.245 (1979) Summary of the Invention [Problem to be solved by the invention]

[0004] The 6DoF content space is characterized by its ability to provide a high sense of realism by simulating the sound arriving at the listening position from an audio object arbitrarily arranged in the 4π space or its acoustic characteristics. The 4π space means the entire space in which the solid angle centered on the listening position or each sound source position is 4π. On the other hand, some audio information presented in the 6DoF content space is not intended to improve the sense of realism. Examples of such audio information include emergency news provided during the broadcast of the main broadcast program, and guidance information (instructions) for device operation. When such audio information (sometimes referred to as "wildcard audio objects" in this application) is reproduced with a specific position in the virtual space as the target position, a situation may occur in which important information does not reach the listener due to distance attenuation depending on the listening position, obstruction by walls, and masking effects by other audio objects (e.g., essential audio objects).

[0005] The embodiments of the present application have been made to solve the above-mentioned problems, and one of the objectives of the present application is to provide an audio processing device and a program therefor that can reliably transmit audio information outside of the content to a listener using an object-based audio system. [Means for solving the problem]

[0006] [1] One aspect of this embodiment is an audio processing device that includes a target direction setting unit that sets a first type target direction, which is a target direction based on a listening position for a first type audio object included in content, and a plurality of second type target directions based on the listening position for a second type audio object, which is an audio object not included in the content, and sets a sum of the directional vectors of the second type target directions to zero.When the minimum value of the angle between the first type target direction and the second type target direction is smaller than a predetermined angle reference value, the target direction setting unit updates the second type target direction so that the minimum value becomes larger. According to the configuration of [1], the first type audio object and the second type audio object are distinguished in terms of the listener's spatial perception, and the listener can lose spatial perception of a specific target direction of the second type audio object by making the sum of the directional vectors of the second type target direction zero. By maintaining spatial perception of the first type target direction for the first type audio object, the listener can easily identify the second type target direction in which spatial perception is lost, so that the listener can reliably transmit audio information to the listener by the second type audio object outside the content.

[0007] [2] One aspect of this embodiment is the above-mentioned audio processing device, wherein the target direction setting unit may rotate the second type target direction around the listening position by a common rotation angle when the minimum value is smaller than the angle reference value. According to the configuration of [2], the sum of the direction vectors of the second type target directions can be maintained at zero by rotating the multiple second type target directions at a common rotation angle. Therefore, it is possible to maintain a state in which the spatial perception of the specific target direction of the second type audio object is lost.

[0008] [3] One aspect of this embodiment is the above-mentioned voice processing device, wherein when the number of first type voice objects being emitted is one, the target direction setting unit may determine a rotation angle of the second type target direction so that, when the minimum value of the angle between the first type target direction and the second type target direction is smaller than the angle reference value, the minimum value is greater than or equal to the angle reference value. According to the configuration of [3], the second type audio object can be distinguished from the first type audio object in terms of spatial perception, and spatial perception of a specific target direction of the second type audio object can be lost. Therefore, the second type target direction can be easily identified.

[0009] [4] One aspect of this embodiment is the above-mentioned audio processing device, wherein when the number of the first type audio objects is multiple, the target direction setting unit may determine a rotation angle of the second type target direction so that, when the minimum value of the angle between the first type target direction and the second type target direction is smaller than the angle reference value, the minimum value becomes larger. According to the configuration of [4], it is possible to distinguish the first type audio objects as much as possible in terms of spatial perception, and to lose spatial perception of a specific target direction of the second type audio object. Therefore, it is possible to easily identify the second type target direction.

[0010] [5] One aspect of the present embodiment may be the above-described audio processing device, further comprising a position information setting unit that sets the listening position or a target position of the first type audio object. According to the configuration of [5], the second type target direction is updated according to the set listening position or the target position of the first type audio object. Therefore, even if the listening position or the target position of the first type audio object changes, audio information by the second type audio object can be transmitted to the listener more reliably.

[0011] [6] One aspect of the present embodiment may be a program for causing a computer to function as the above-described voice processing device. According to the configuration of [6], the first type audio object and the second type audio object are distinguished in terms of the listener's spatial perception, and the listener can lose spatial perception of a specific target direction of the second type audio object by making the sum of the directional vectors of the second type target direction zero. By maintaining spatial perception of the first type target direction for the first type audio object, the listener can easily identify the second type target direction in which spatial perception is lost, so that the listener can reliably transmit audio information to the listener by the second type audio object outside the content. Effect of the Invention

[0012] According to this embodiment, audio information outside the content can be reliably transmitted to the listener. [Brief description of the drawings]

[0013] [Figure 1] 1 is a schematic block diagram illustrating an overview of a voice processing system according to an embodiment of the present invention. [Diagram 2]FIG. 11 is an explanatory diagram showing an example of setting a target direction of a voice object. [Diagram 3] 1 is a schematic block diagram illustrating an example of a functional configuration of a sound processing device according to an embodiment of the present invention. [Figure 4] FIG. 4 is a schematic block diagram showing another example of the functional configuration of the audio processing device according to the present embodiment. [Diagram 5] 10 is a flowchart illustrating an example of a target direction setting process according to the embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0014] (overview) First, an overview of an embodiment of the present application will be described with reference to the drawings. FIG. 1 is a schematic block diagram illustrating an overview of a sound processing system 1 according to the present embodiment. The sound processing system 1 includes a sound processing device 10, a playback processing unit 130, and a playback device 20. The sound processing system 1 is an example of an object-based audio system. In the following description, the sound processing system 1 is mainly configured to be capable of playing 6DoF sound content in a three-dimensional space.

[0015] The audio processing device 10 acquires audio content. The audio content includes an audio signal and metadata for each audio object. The audio content may include at least one audio object as a mandatory audio object. The mandatory audio object is an audio object whose playback characteristics are variable depending on a user's operation or the environment. The variable playback characteristics may include volume. The mandatory audio object is also called a priority audio object. For example, the mandatory audio object may be applied to the voice of the lines of the performers in a movie or a drama, or the commentary voice in a sports broadcast or a news program. The mandatory audio object is identified by a mandatory flag included in the metadata.

[0016] The audio processing device 10 acquires audio content (sometimes referred to as "secondary audio content" in this application) separate from the primary audio content. The secondary audio content may include a wildcard audio object. Examples of the wildcard audio object include emergency broadcast audio, audio guidance for operating various devices, and the like. The wildcard audio object is identified by a wildcard flag included in the metadata.

[0017] The audio processing device 10 sets spatial information for perceiving a sound image for each audio object, and outputs the determined spatial information and audio signal to the reproduction processing unit 130. More specifically, the audio processing device 10 is capable of setting a target position and user position information for each required audio object. The number of required audio objects can be one or more than two. The user position information can include one or both of the position (i.e., listening position) and the orientation (i.e., listening direction) of the listener as a user. The audio processing device 10 notifies the reproduction processing unit 130 of a target position based on the listening position for the required audio object. The target position corresponds to an element of spatial information including a distance from an origin and a target direction. The audio processing device 10 determines a target direction as spatial information of the wildcard audio object based on the target direction of the required audio object using a method described later. The number of wildcard audio objects is usually one. However, the number of target directions for the wildcard object is set to a predetermined integer of two or more. The audio processing device 10 notifies the playback processing unit 130 of the target direction of the wildcard audio object.

[0018] Based on the spatial information and audio signal of the audio object input from the audio processing device 10, the playback processing unit 130 generates an object-specific audio signal to be used for playback according to the spatial information. The playback processing unit 130 adds acoustic characteristics corresponding to a target direction to the input audio signal. For an audio object to which information of a target position is given, the playback processing unit 130 adjusts a gain corresponding to the distance to which the acoustic characteristics are added. In adjusting the gain, a predetermined distance attenuation model (e.g., the inverse square law) is used. The playback processing unit 130 outputs a playback signal based on the generated object-specific audio signal to the playback device 20. Thus, the playback device 20 can present the audio content and the non-content audio so that they are perceived from different target positions by the playback signal acquired from the playback processing unit 130.

[0019] A method corresponding to the playback method used in the playback device 20 is used as a process for adding the acoustic characteristics. When the playback method is multi-channel audio, playback processing information indicating a reference direction of the playback sound source for each playback channel constituting the playback device 20 and a gain for each channel corresponding to the target direction is set in the playback processing unit 130. By referring to the playback processing information, the playback processing unit 130 can specify two or more playback channels that spatially include a target direction for an audio object and specify a gain corresponding to the reference direction for each playback channel. The playback processing unit 130 can generate a multi-channel object-specific audio signal whose amplitude is adjusted by multiplying the acquired audio signal by the gain specified for each playback channel. The playback processing unit 130 can generate a signal obtained by adding the object-specific audio signal between audio objects for each playback channel as a playback signal.

[0020] When the playback method is binaural sound, playback processing information indicating a head-related transfer function corresponding to a target direction for each playback channel related to the playback device 20 is set in advance in the playback processing unit 130. By referring to the playback processing information, the playback processing unit 130 can identify a head-related transfer function corresponding to a target direction for an audio object for each playback channel, and convolve the identified head-related transfer function with the input audio signal to generate a two-channel object-specific audio signal. The playback processing unit 130 can generate a signal obtained by adding the generated object-specific audio signals between objects for each playback channel as a playback signal.

[0021] Next, an example of setting the target direction for each audio object will be described. FIG. 2 is an explanatory diagram showing an example of setting the target direction of an audio object. FIG. 2 illustrates three essential audio objects EO01, EO02, and EO03 included in 6DoF audio content. The target positions of the essential audio objects may be designated in advance by metadata, or the audio processing device 10 may set the target positions of the essential audio objects in response to an operation. The target direction of the essential audio object EO01 is set in a direction diagonally upward and behind the listener. The target direction of the essential audio object EO02 is set in a direction diagonally upward and in front of the listener. The target direction of the essential audio object EO03 is set in a direction diagonally downward and behind the listener. The essential audio objects are applied as main components of the audio content. For example, speech (also called dialogue) of a commentator, a presenter, a guest, a character, etc., singing, musical instrument sounds, sound effects, etc. may be applied as essential audio objects.

[0022] The voice processing device 10 sets a plurality of target directions for the wildcard voice object such that the sum of the individual directional vectors is zero. The directional vector is a vector indicating a target direction whose magnitude is normalized to 1. In the example of FIG. 2, six target directions p1 to p16 are set. The end points of the directional vectors indicating the target directions p1 to p6 starting from the origin are distributed on a sphere of radius 1 in a three-dimensional space. The directional vectors indicating the target directions p1 to p6 are expressed in polar coordinates as (1, π / 4, π / 2), (1, 3π / 4, π / 3), (1, 5π / 4, π / 2), (1, 7π / 4, π / 2), (1, 0, 0), and (1, 0, π), respectively. Each of the illustrated polar coordinates includes a distance r from the origin, an azimuth angle φ in the horizontal plane, and a vertex angle θ, in that order.

[0023] For the azimuth angle, 0 degrees is to the right of the origin. For the vertex angle, 0 degrees is the normal direction to the horizontal plane. That is, the target direction p1 is the direction diagonally forward to the right of the listener in the horizontal plane, the target direction p2 is the direction diagonally forward to the left of the listener in the horizontal plane, and the target direction p3 is the direction diagonally behind the left of the listener in the horizontal plane. The target direction p4 is the direction diagonally behind the right of the listener in the horizontal plane, the target direction p5 is the direction directly above the listener, and the target direction p6 is the direction directly below the listener.

[0024] The voice processing device 10 simultaneously changes a plurality of target directions related to the wildcard voice object so that the minimum value of the angle between the target direction of the essential voice object and the target direction of the wildcard voice object becomes larger. More preferably, the voice processing system 1 sets the minimum value of the angle to a predetermined angle reference value or more, and rotates the target direction of the wildcard voice object by a common rotation angle so that the directional distribution between the target directions is maintained. Thus, the minimum angle between the target directions of the essential voice object and the wildcard voice object increases while keeping the sum of the target direction vectors for each changed target direction at the origin. The angle reference value may be any angle that can reliably distinguish the directions in which a plurality of voice objects are perceived. For example, π / 9, which is an angle that can reliably distinguish the voices of each of a plurality of objects appearing in a video while presenting the video, may be applied as the angle reference value (see Non-Patent Document 8).

[0025] According to the reproduction processing unit 130, acoustic characteristics corresponding to the set target direction are added to the audio signal of the essential audio object. Meanwhile, acoustic characteristics corresponding to each of the multiple target directions are added to the audio signal of the wildcard audio object. The reproduction device 20 presents a sound based on a reproduction signal obtained by mixing the audio signals to which the acoustic characteristics have been added. For the essential audio objects, the listener is provided with the perception of the individually set target direction, whereas for the wildcard audio object mixed among the multiple target directions, the listener loses the perception of the individual target direction. Therefore, the content of the wildcard audio object can be reliably conveyed to the listener in distinction from the essential audio object with the perception of the target direction maintained.

[0026] Next, an example of the functional configuration of the voice processing device 10 according to the present embodiment will be described. Fig. 3 is a schematic block diagram showing an example of the functional configuration of the voice processing device 10 according to the present embodiment. However, in the example of Fig. 3, it is assumed that the number of required voice objects is one. The voice processing device 10 includes a position information setting unit 110 and a target direction setting unit 120.

[0027] The position information setting unit 110 sets user position information relating to the listening position of the user. When setting the user position information, the position information setting unit 110 may display, for example, a setting screen showing a listening space including the user's head and its surroundings on a display unit (not shown). The display unit may be, for example, any of a display monitor, a display panel, and the like. The position information setting unit 110 may determine, as a listening position, a position corresponding to a position in the listening space designated by an operation signal input from an input unit (not shown), and may determine, as a listening direction, a direction of the head designated by the operation signal. The input unit may be, for example, any of a mouse, a touch sensor, a keyboard, a button, a dial, a knob, and the like. The position information setting unit 110 may obtain a position detection signal indicating the detected position from a position sensor (not shown) that detects the position of the user. The position information setting unit 110 may obtain a direction detection signal indicating the detected direction from a direction sensor (not shown) that detects the direction of the user's head.

[0028] The target direction setting unit 120 sets a target direction of the voice object. The target direction setting unit 120 includes a reference distance conversion unit 121, a minimum angle interval selection unit 122, a reference rotation vector derivation unit 123, a rotation angle determination unit 124, and a voice object rotation unit 125.

[0029] The reference distance conversion unit 121 converts the target position of the essential voice object to a position a predetermined reference distance away from the listening position without changing the direction from the listening position. The reference distance conversion unit 121 acquires a required audio object and receives user position information from the position information setting unit 110. The reference distance conversion unit 121 converts the target position extracted from the metadata of the required audio object into a relative position s i (r i ,φ i ,θ iThe reference distance conversion unit 121 converts the relative position obtained by the conversion into a position that is a predetermined reference distance away from the origin without changing the direction from the origin (for example, if the reference distance is 1, then (1,φ i ,θ i ) is normalized to the target position s i That is, the reference distance conversion unit 121 determines the intersection point obtained by projecting the relative position of the essential voice object onto the sphere of the reference distance from the origin as the normalization target position. i ,θ i ) corresponds to the target direction of the required voice object i. The reference distance conversion unit 121 outputs target direction information indicating the target direction to the minimum angle interval selection unit 122 .

[0030] The minimum angle interval selection unit 122 receives target direction information from the reference distance conversion unit 121. In addition, the target direction setting unit 120 presets initial values ​​of a plurality of target directions of the wildcard voice object. When a voice signal of the wildcard voice object is input from the outside, the minimum angle interval selection unit 122 selects a plurality of target directions p of the wildcard voice object. j (j is an integer between 1 and J, and J is the number of wildcard target directions), the target direction p that forms the smallest angle with the target direction of the required voice object j In the following description, the selected target direction may be referred to as the "minimum angle target direction", and the angle between the minimum angle target direction and the target direction of the required voice object may be referred to as the "minimum angle". The minimum angle interval selection unit 122 outputs wildcard voice object information indicating the minimum angle target direction to the reference rotation vector derivation unit 123.

[0031] The reference rotation vector derivation unit 123 identifies the minimum angle target direction indicated by the wildcard audio object information input from the minimum angle interval selection unit 122 and the target direction indicated by the target direction information added to the audio signal of the required audio object. The reference rotation vector derivation unit 123 calculates the minimum angle target direction p j' and the target direction s indicated by the target direction information added to the speech signal of the essential speech object. i ' and the angle ξ si’,pj’ It is determined whether the angle ξ is smaller than a predetermined angle reference value. si’,pj’ The angle reference value corresponds to the above-mentioned minimum angle.

[0032] angle ξ si’,pj’ is smaller than the predetermined angle reference value, the reference rotation vector derivation unit 123 determines that the target direction of the wildcard voice object is to be rotated. In this case, the reference rotation vector derivation unit 123 determines that the target direction s of the mandatory voice object is to be rotated. i The rotation vector indicating the rotation around the origin from the minimum angle ' to the target direction is the reference rotation vector R si’,pj’ (·). This rotation is defined as the direction vector s i ' and the direction vector p j The rotation is performed in the plane spanned by the reference rotation vector R si’,pj’ (·) is, for example, the direction vector s i ' and the direction vector p j Diff of 's i '-p j The reference rotation vector derivation unit 123 calculates the reference rotation vector R si’,pj’ (·) is output to the rotation angle determination unit 124.

[0033] angle ξ si’,pj’ is equal to or greater than a predetermined angle reference value, the reference rotation vector derivation unit 123 determines not to rotate the target direction of the wildcard voice object. In this case, the reference rotation vector derivation unit 123 determines that the reference rotation vector R si’,pj’ (·), the rotation angle determination unit 124 does not need to calculate the rotation angle. The voice object rotation unit 125 does not perform a rotation operation on the target direction of the wildcard voice object, but adopts the initial values. In this case, the reference rotation vector derivation unit 123 does not calculate the reference rotation vector R si’,pj’ (·) may be defined as zero, which also effectively disables rotation of the wildcard audio object to the target direction.

[0034] The rotation angle determination unit 124 determines a rotation angle for rotating the target direction of the wildcard voice object so that the angle between the minimum angle target direction and the target direction of the essential voice object is changed from the minimum angle to an angle reference value. si’,pj’ is the target direction s of the required speech object i ' and the minimum angle target direction p j ' and the angle ξ si’,pj’ The rotation angle ξ si’,pj’ is the target direction s of the required speech object as shown in equation (1). i ' indicates the direction vector s i + and the minimum angle target direction p j ' indicates the direction vector p j + + is divided by the absolute value of each, and the angle corresponds to the cosine value. In equation (1), arccos(...) is the inverse cosine function of.... Direction vector s i +, p j + is expressed in an orthogonal coordinate system. The rotation angle determination unit 124 calculates the rotation angle information indicating the calculated rotation angle as the reference rotation vector R si’,pj’ The voice object is then associated with (·) and output to the voice object rotation unit 125.

[0035]

number

[0036] The voice object rotation unit 125 determines the rotation angle ξ input from the rotation angle determination unit 124. si’,pj’ and the reference rotation vector R si’,pj’ Based on (·), the target direction p of the wildcard speech object j Rotate around the origin. Rotation angle ξ si’,pj’ is greater than 0 and smaller than the angle reference value, the audio object rotator 125 rotates the reference rotation vector R si’,pj’ The common rotation vector expressed by (·) using a coefficient k that indicates the number of times the reference rotation vector acts k R si’,pj’Using (·), multiple target directions p j The target direction p after the rotation operation is j " is expressed by equation (2). The voice object rotation unit 125 changes the angle reference value ξ to a rotation angle ξ so as to satisfy the condition shown in the formula (3). si’,pj’ The coefficient k can be calculated by dividing by . The coefficient k is a positive real number greater than 1.

[0037]

number

[0038]

number

[0039] In addition, the rotation angle ξ si’,pj’ When the target direction p of the wildcard voice object after rotation is 0, the relationship shown in the formulas (2) and (3) cannot be used. j " and the target direction s of the required speech object i ' and the angle ξ si’,pj’ The rotation vector R0(·) is set in advance so that the rotation angle ξ is equal to or greater than the reference angle ξ. si’,pj’ becomes 0, the voice object rotation unit 125 rotates the voice object in the target direction p using a preset rotation vector R0(·). j Rotate each of the following.

[0040] When the voice object rotation unit 125 rotates the target direction of the wildcard voice object, the target direction p j The metadata including the target direction information indicating the wildcard audio object is added to the audio signal of the wildcard audio object and output to the playback processing unit 130. When it is determined that the target direction of the wildcard voice object is not to be rotated, the voice object rotation unit 125 rotates the target direction p jThe metadata including the target direction information indicating the wildcard audio object is added to the audio signal of the wildcard audio object and output to the playback processing unit 130.

[0041] Fig. 4 is a schematic block diagram showing another example of the functional configuration of the voice processing device 10 according to the present embodiment. In the example of Fig. 4, a case where the number of required voice objects is multiple is taken as an example. The explanation of Fig. 4 will mainly focus on the differences from the example of Fig. 3, and the explanation of Fig. 3 will be used unless otherwise stated. The target direction setting unit 120 includes a reference distance conversion unit 121 , a minimum angle interval selection unit 122 , a reference rotation vector derivation unit 123 , a voice object rotation unit 125 , and a rotation amount evaluation unit 126 .

[0042] Audio signals for a plurality of essential audio objects can be externally input to the reference distance conversion unit 121. The reference distance conversion unit 121 converts the target position for each essential audio object into a relative position with the listening position as the origin, and determines a normalized target position whose distance from the origin is the reference distance. The minimum angle interval selection unit 122 specifies an angle between the target direction of the required voice object and the target direction of a predefined wildcard voice object for each pair. The minimum angle interval selection unit 122 selects a minimum angle min{ξ si’,pj} is the target direction of the required speech object s m and the target direction of the wildcard speech object p n In the following description, the selected set may be referred to as the "minimum angle set."

[0043] The reference rotation vector derivation unit 123 determines whether or not to rotate the target direction of the wildcard voice object depending on whether or not the minimum angle is smaller than a predetermined angle reference value. When rotating the target direction of the wildcard voice object, the reference rotation vector derivation unit 123 rotates the target direction s of the required voice object related to the minimum angle set. m A rotation vector indicating the rotation around the origin from the wildcard audio object to the reference rotation vector Rsm,pn It is calculated as (·). When the target direction of the wildcard voice object is not rotated, the reference rotation vector derivation unit 123 determines the reference rotation vector R sm,pn (·), and the rotation amount evaluation unit 126 does not calculate a coefficient k0 (described later) based on the evaluation function (k). Then, the voice object rotation unit 125 does not perform a rotation operation on the target direction of the wildcard voice object, but adopts their initial values. In this case, the reference rotation vector derivation unit 123 adopts their initial values. In this case, the reference rotation vector derivation unit 123 calculates the reference rotation vector R si’,pj’ (·) may be defined as zero, which also effectively disables rotation of the wildcard audio object to the target direction.

[0044] The rotation amount evaluation unit 126 calculates the reference rotation vector R sm,pn Based on (·), the target direction s of each required speech object i ' and the target direction of the wildcard speech object after rotation p j The angle ξ si’,pj” Calculate the evaluation function E(k) for evaluating the validity of each angle ξ si’,pj” Any function that gives a function value that monotonically decreases as ξ increases may be used. k is a real number greater than 1. The rotation amount evaluation unit 126 may use, for example, an evaluation function E(k) shown in equation (4). According to equation (4), any angle ξ si’,pj” As k approaches 0, the evaluation function E(k) diverges toward infinity.

[0045]

number

[0046] Then, the rotation amount evaluation unit 126 calculates the coefficient k and the minimum angle ξ as shown in the formula (5). sm’,pn Product kξ sm’,pnis equal to or smaller than the reference angle value ξ, the coefficient k that minimizes the evaluation function E(k) is set as the coefficient k0. Therefore, the largest possible angle ξ for the entire system is si’,pj” The coefficient k0 is determined so that the rotation amount evaluation unit 126 calculates the coefficient k0 and the reference rotation vector R sm’,pn (·) is output to the voice object rotation unit 125.

[0047]

number

[0048] The voice object rotation unit 125 calculates the coefficient k0 input from the rotation amount evaluation unit 126 and the reference rotation vector R sm,pn Based on (·), the target direction p of the wildcard speech object j Rotate around the origin. minimum angle ξ sm’,pn If R is greater than 0 and smaller than the angle reference value, the audio object rotator 125 sm’,pn The common rotation vector expressed in (·) with coefficient k0 k0 R sm,pn Using (·), the target direction p of the wildcard speech object j The target direction p after the rotation operation is j " is expressed by equation (6).

[0049]

number

[0050] In addition, the minimum angle ξ sm’,pn becomes 0, the voice object rotation unit 125 rotates the voice object in the target direction p using a preset rotation vector R0(·). j The rotation vector R0(·) is the target direction p of the wildcard voice object after rotation. j " and the target direction s of the required speech object i ' and the smallest angle ξ sm’,pnis previously determined to be at least greater than 0, and preferably equal to or greater than the reference angle value ξ.

[0051] As the initial value of the target direction of the wildcard audio object, it may be desirable to avoid the front direction (φ, θ) = (π / 2, π / 2) of the listening position or directions within a predetermined angle reference value (e.g., π / 9) from that direction, as exemplified in Fig. 2. This is because, in general, listeners tend to pay attention to sounds perceived as being in front, and producers tend to provide audio objects whose target direction is the front. In addition, it is desirable for the interval between adjacent target directions to be equal to or greater than a predetermined angle reference value.

[0052] In each of the above configuration examples, the target direction setting unit 120 may receive the mandatory audio objects and the wildcard audio objects from a broadcasting facility via a broadcast transmission path, or from another device via a communication network. The target direction setting unit 120 may read out the mandatory audio objects stored in advance in a storage medium provided in or connected to the audio processing device 10. The mandatory audio object is included in the main content selected directly in response to the operation, or in the main content provided in the broadcast channel or communication directly instructed in response to the operation. The wildcard audio object is usually provided asynchronously with the mandatory audio object. The wildcard audio object may be provided in response to the operation or may be provided regardless of the operation, depending on the information transmitted.

[0053] 2 illustrates an example in which the number of target directions of the wildcard voice object is six, but this is not limited as long as the above conditions are satisfied and the sum of the direction vectors for each target direction can be set to 0. The number of target directions may be 2 to 5, or 7 or more. When the number of target directions is two, the initial values ​​of the target directions p1 and p2 can be set to, for example, (1, π / 2, π / 2) and (1, 3π / 2, π / 2). When the number of target directions is four, the initial values ​​of the target directions p1 to p4 can be set to, for example, (1, π / 4, π / 2), (1, 3π / 4, π / 2), (1, 5π / 4, π / 2), and (1, 7π / 4, π / 2). In these examples, on the xy plane the target directions are positioned on the straight lines y=±x. Furthermore, the multiple target directions may be distributed on a single straight line forming a one-dimensional space, for example, on either the straight line y=+x or y=-x, and may not be distributed in other regions. Therefore, placement in the front direction of the listening position and within the range of the angle reference value from that direction is avoided.

[0054] Next, an example of a process for setting a target direction of a wildcard voice object according to this embodiment will be described. Fig. 5 is a flowchart showing an example of a process for setting a target direction according to this embodiment. The process in Fig. 5 is an example in which the number of required voice objects is one, similar to Fig. 3.

[0055] (Step S102) The target direction setting unit 120 acquires a required voice object. (Step S104) The target direction setting unit 120 acquires a wildcard voice object. (Step S106) The position information setting unit 110 acquires user position information indicating the listening position.

[0056] (Step S108) The target direction setting unit 120 converts the target position of the essential sound object to a position whose distance from the target position is the reference distance, without changing the direction from the listening position. (Step S110) The target direction setting unit 120 selects the target direction that forms the smallest angle with the target direction of the required voice object from among multiple target directions of the pre-set wildcard voice object as the minimum angle target direction, and identifies that angle as the minimum angle.

[0057] (Step S112) The target direction setting unit 120 judges whether the specified minimum angle is smaller than a predetermined angle reference value. If it is judged that the minimum angle is smaller than the angle reference value (step S112 YES), the process proceeds to step S114. If it is judged that the minimum angle is equal to or larger than the angle reference value (step S112 NO), the target direction setting unit 120 notifies the playback processing unit 130 of the initial value of the target direction set in advance of the wildcard voice object, and the process proceeds to step S120.

[0058] (Step S114) The target direction setting unit 120 calculates a reference rotation vector indicating the rotation from the target direction of the essential voice object to the minimum angle target direction. (Step S116) The target direction setting unit 120 determines a rotation angle based on the target direction of the essential voice object and the minimum angle target direction. (Step S118) The target direction setting unit 120 rotates the initial value of the target direction of the wildcard voice object using the calculated reference rotation vector and rotation angle. The target direction setting unit 120 notifies the playback processing unit 130 of the target direction of the wildcard voice object after rotation.

[0059] (Step S120) The playback processing unit 130 performs playback processing on the audio signal of the audio object according to the playback method of the playback device 20 based on the target direction set for each audio object, to generate an object-specific audio signal. When performing playback processing on a required audio object, the playback processing unit 130 refers to the target direction set for the required audio object. When performing playback processing on a wildcard audio object, the playback processing unit 130 uses each of the multiple target directions notified by the target direction setting unit 120 for the audio signal to generate object-specific audio signals in the number of objects. The playback processing unit 130 mixes the object-specific audio signals generated between the audio object and the target directions for each playback channel to generate a playback signal, and outputs the playback signal to the playback device 20. The playback device 20 presents a playback sound according to the playback signal acquired from the playback processing unit 130. Thereafter, the processing of FIG. 5 ends.

[0060] When there are a plurality of required voice objects, the process of FIG. 5 can be modified as follows. The process of step S108 is performed for each of the required voice objects. In step S110, the target direction setting unit 120 determines, for each pair of the target direction of the essential voice object and the target direction of the predefined wildcard voice object, the minimum angle between them. The target direction setting unit 120 identifies, as a minimum angle set, the pair of the target direction of the essential voice object and the target direction of the wildcard voice object that gives the determined minimum angle.

[0061] In step S114, the target direction setting unit 120 determines a reference rotation vector indicating a rotation around the origin from the target direction of the required voice object associated with the minimum angle set to the wildcard voice object. In step S116, the target direction setting unit 120 searches for a coefficient k0 that minimizes a predetermined evaluation function indicating the validity of the angle between the target direction of each required voice object and the target direction of the rotated wildcard voice object based on the reference rotation vector, under the condition that the product of the coefficient k and the minimum angle is less than or equal to a reference angle value. In step S118, the target direction setting unit 120 rotates the target directions of the wildcard voice objects by a common rotation angle based on the coefficient k0 and the reference rotation vector.

[0062] In the above description, it is assumed that the target position of each required voice object is described in its metadata and notified, but this is not limited to the above. The position information setting unit 110 may set object position information indicating the target position of the required voice object set by the target direction setting unit 120. The position information setting unit 110 can display an icon indicating the required voice object in the listening space on the setting screen, and determine a position corresponding to a position in the listening space indicated by an operation signal as the target position of the required voice object. The position information setting unit 110 outputs the set object position information to the target direction setting unit 120. The target direction setting unit 120 may use the target position indicated by the object position information instead of the target position previously set for the required voice object. Similarly, when a target position is not set for the required voice object, the target direction setting unit 120 may use the target position indicated by the object position information. In that case, the target direction of the wild voice object is set based on the target position arbitrarily set according to the operation.

[0063] In the above description, the target direction setting unit 120 uses the reference rotation vector calculated by the reference rotation vector derivation unit 123 when the voice object rotation unit 125 rotates the initial value of the direction vector indicating the target direction of the wildcard voice object, but the present invention is not limited to this. The voice object rotation unit 125 may rotate the target direction related to each direction vector by multiplying a rotation matrix by the initial value of each direction vector instead of the reference rotation vector. The target direction setting unit 120 may search for a rotation matrix in which the minimum angle calculated by the target direction after rotation becomes the angle reference value. If such a rotation matrix is ​​not determined, the target direction setting unit 120 may calculate a rotation matrix so that the minimum angle becomes larger under the target direction of the given required object (maximization). The voice object rotation unit 125 may multiply the calculated rotation matrix by the initial value of each direction vector to rotate each target direction.

[0064] The voice object rotation unit 125 may preset target directions of a plurality of sets of wildcard voice objects. The target direction for each set is preset so that the angles between the target directions of different sets are equal to or greater than the angle reference value. Therefore, when it is determined that the minimum angle for a certain set is smaller than the angle reference value, the voice object rotation unit 125 may search for another set whose minimum angle is equal to or greater than the angle reference value. When there is no other set whose minimum angle is equal to or greater than the angle reference value, the voice object rotation unit 125 may select another set whose minimum angle is the largest. The voice object rotation unit 125 adopts the target direction of the wildcard voice object related to the determined set.

[0065] In the above description, the audio processing system 1 is mainly assumed to be capable of reproducing 6DoF audio content and to handle listening positions distributed in a three-dimensional space and target positions for each audio object, but this is not limited thereto. The audio processing system 1 may be capable of reproducing at least 4DoF audio content and to handle listening positions distributed in a two-dimensional space and target positions for each audio object. In this case, the polar coordinates in the two-dimensional space are equivalent to the polar coordinates in the three-dimensional space with the vertex angle omitted. In the audio processing system 1, the audio processing device 10 may be configured to include a playback processing unit 130. The audio processing device 10 may be configured to further include a playback device.

[0066] As described above, the audio processing device 10 according to the present embodiment includes a target direction setting unit 120 that sets a first type target direction, which is a target direction based on a listening position for a first type audio object (e.g., a required audio object) included in the content, and a plurality of second type target directions (e.g., target directions of wildcard audio objects) based on listening positions for a second type audio object (e.g., a wildcard audio object) that is an audio object not included in the content. The sum of the direction vectors of the second type target directions is set to zero. When the minimum value (e.g., minimum angle) of the angle between the first type target direction and the second type target direction is smaller than a predetermined angle reference value, the target direction setting unit 120 updates the second type target direction so that the minimum value becomes larger. According to this configuration, the first type audio object and the second type audio object are distinguished in terms of the listener's spatial perception, and the listener can lose spatial perception of a specific target direction of the second type audio object by making the sum of the directional vectors of the second type target direction zero. By maintaining spatial perception of the first type target direction for the first type audio object, the listener can easily identify the second type target direction in which spatial perception is lost, so that the listener can reliably transmit audio information to the listener by the second type audio object outside the content.

[0067] When the minimum value is smaller than the angle reference value, the target direction setting section 120 may rotate the second type target direction around the listening position by a common rotation angle. According to this configuration, the plurality of second type target directions rotate at a common rotation angle, so that the sum of the direction vectors of the second type target directions can be maintained at zero, thereby making it possible to maintain a state in which the spatial perception of the specific target direction of the second type audio object is lost.

[0068] When the number of first type voice objects is one, the target direction setting unit 120 may determine the rotation angle of the second type target direction so that when the minimum value of the angle between the first type target direction and the second type target direction is smaller than the angle reference value, the minimum value is equal to or greater than the angle reference value. According to this configuration, the second type audio object can be distinguished from the first type audio object in terms of spatial perception, and spatial perception of a specific target direction of the second type audio object can be lost. Therefore, the second type target direction can be easily identified.

[0069] When there are multiple first type voice objects, the target direction setting unit 120 may determine the rotation angle of the second type target direction so that the minimum value of the angle between the first type target direction and the second type target direction becomes larger when the minimum value of the angle is smaller than the angle reference value. According to this configuration, it is possible to distinguish the first type sound objects as much as possible in terms of spatial perception, and to lose spatial perception of a specific target direction of the second type sound object, so that the second type target direction can be easily identified.

[0070] The sound processing device 10 according to this embodiment may include a position information setting unit 110 that sets the listening position or the target position of the first type sound object. According to this configuration, the second type target direction is updated according to the listening position or the target position of the first type audio object set by the position information setting unit 110. Therefore, even if the listening position or the target position of the first type audio object changes, audio information by the second type audio object can be transmitted to the listener more reliably.

[0071] In addition, all or a part of the above-mentioned voice processing device 10, for example, the position information setting unit 110 and the target direction setting unit 120, may be realized by a computer. The function may be realized by recording a program for realizing the function in a computer-readable recording medium, reading the program recorded in the recording medium into a processor of a computer system, and executing a process instructed by an instruction described in the program. In addition, the "computer system" here refers to a computer system built into the voice processing device 10, and includes hardware such as an OS (Operating System) and peripheral devices. In addition, the "computer-readable recording medium" refers to a portable medium such as a flexible disk, an optical magnetic disk, a ROM, a CD-ROM, and a storage device such as a hard disk built into a computer system. Furthermore, the "computer-readable recording medium" may include a medium that dynamically holds a program for a short time, such as a communication line when a program is transmitted via a network such as the Internet or a communication line such as a telephone line, and a medium that holds a program for a certain period of time, such as a volatile memory inside a computer system that is a server or client in that case. In addition, the above-mentioned program may be a program for realizing a part of the above-mentioned function, and may further be a program that can realize the above-mentioned function in combination with a program already recorded in the computer system.

[0072] In addition, a part or the whole of the voice processing device 10 in the above-mentioned embodiment may be realized as an integrated circuit such as an LSI (Large Scale Integration). Each functional block of the voice processing device 10 may be individually processed, or a part or the whole may be integrated and processed. The integrated circuit method is not limited to LSI, and may be realized by a dedicated circuit or a general-purpose processor. In addition, when an integrated circuit technology that replaces LSI appears due to the progress of semiconductor technology, an integrated circuit based on that technology may be used.

[0073] Although one embodiment of the present invention has been described in detail above with reference to the drawings, the specific configuration is not limited to the above, and various design changes, etc. are possible within the scope that does not deviate from the gist of the present invention. [Explanation of symbols]

[0074] 1...sound processing system, 10...sound processing device, 112...coordinate conversion unit, 114...distance attenuation characteristic generation unit, 116...conversion filter generation unit, 118...filter processing unit, 1142...position correction processing unit, 1144...head diffraction processing unit, 1146...distance attenuation characteristic calculation unit

Claims

1. a first type target direction that is a target direction based on a listening position for a first type sound object included in the content; a target direction setting unit that sets a plurality of second target directions based on the listening position of a second type sound object that is a sound object not included in the content; The sum of the direction vectors of the second type target direction is set to zero, The target direction setting unit is When the minimum value of the angle between the first type target direction and the second type target direction is smaller than a predetermined angle reference value, the second type target direction is updated so that the minimum value becomes larger. Audio processing device.

2. The target direction setting unit is When the minimum value is smaller than the angle reference value, the second type target direction is rotated around the listening position by a common rotation angle. The audio processing device according to claim 1 .

3. When the number of the first type voice object is one, The target direction setting unit is When a minimum value of an angle between the first type target direction and the second type target direction is smaller than the angle reference value, a rotation angle of the second type target direction is determined so that the minimum value is equal to or larger than the angle reference value. The audio processing device according to claim 1 .

4. When the number of the first type voice objects is more than one, The target direction setting unit is When a minimum value of an angle between the first type target direction and the second type target direction is smaller than the angle reference value, a rotation angle of the second type target direction is determined so that the minimum value becomes larger. The audio processing device according to claim 1 .

5. a position information setting unit that sets the listening position or a target position of the first type sound object; The audio processing device according to claim 1 .

6. A program for causing a computer to function as the voice processing device according to claim 1.

Citation Information

Patent Citations

  • ITRBS.2125-0,

  • ITRBS.2076-1,

  • IEC23008-3