Audio reproduction methods and sound reproduction systems
By using multiple rows of audio transducers and metadata processing, the problem of sound and visual cues being disconnected in existing audio systems has been solved, achieving a realistic perception of audio near visual cues and a position-insensitive listening effect.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2011-03-17
- Publication Date
- 2026-04-03
AI Technical Summary
Existing audio systems lack visual cues when localizing sound, resulting in unrealistic sound perception and often requiring a fixed listener position for optimal effect.
By using multiple rows of audio transducers, combined with metadata and signal processing techniques, the origin of the audio signal is determined, and weighting factors are used to present the audio signal, so that sound perception corresponds to visual cues on the video plane, thereby achieving localized audio perception.
It improves the perceived realism of sound near visual cues, reduces dependence on the listener's location, enhances the integration of audio and video scenes, and improves the listening experience.
Smart Images

Figure CN116437283B_ABST
Abstract
Description
[0001] This application is a divisional application of patent application number 202110746811.6, filed on March 17, 2011, entitled "Audio Reproduction Method and Sound Reproduction System". Patent application number 202110746811.6 is a divisional application of patent application number 201810895098.X, filed on March 17, 2011, also entitled "Audio Reproduction Method and Sound Reproduction System". Patent application 0895098.X is a divisional application of patent application No. 201510284811.3, filed on March 17, 2011, entitled "Technology for Localized Audio Sensing". Patent application No. 201510284811.3 is a divisional application of patent application No. 201180015018.3, filed on March 17, 2011, entitled "Technology for Localized Audio Sensing".
[0002] Cross-reference to related applications
[0003] This application claims priority to U.S. Provisional Patent Application No. 61 / 316,579, filed March 23, 2010, the entire contents of which are incorporated herein by reference. Technical Field
[0004] The present invention generally relates to audio reproduction, and more particularly to audio perception in the vicinity of visual cues. Background Technology
[0005] Whether in a residential living room or a theater, high-fidelity sound systems approximate the actual original sound field by utilizing stereo technology. These systems use at least two presentation channels (e.g., left and right channels, surround sound 5.1, 6.1, or 11.1, etc.), typically projected through a symmetrical arrangement of speakers. For example, as shown in Figure 1, a typical surround sound 5.1 system 100 includes: (1) a left front speaker 102, (2) a right front speaker 104, (3) a front center speaker 106 (center channel), (4) a low-frequency speaker 108 (e.g., a subwoofer), (5) a left rear speaker 110 (e.g., left surround), and (6) a right rear speaker 112 (e.g., right surround). In system 100, the front center speaker 106 or a single center channel carries all dialogue and other audio associated with images on the screen.
[0006] However, these systems suffer from drawbacks, particularly in localizing sound in certain directions, and often require a single, fixed listener position for optimal performance (e.g., optimal listening position 114, the focal point between the speakers where the individual hears the audio mix the mixer wants). Many efforts to date to improve this have involved increasing the number of presentation channels. Mixing a large number of channels incurs greater time and cost disadvantages for content creators, yet the resulting perception fails to confine the sound to the vicinity of the visual cues of the sound source. In other words, the sound reproduced from these sound systems is not perceived as emanating from the on-screen video plane, thus lacking a true sense of realism.
[0007] The inventors realized from the above that the technology for localized perceptual audio associated with video images is needed to improve the natural listening experience.
[0008] The methods described in this section are methods that can be sought, but are not necessarily methods that have been previously conceived or sought. Therefore, unless otherwise stated, one should not assume that a method is considered prior art simply because any method described in this section is included in this section. Similarly, unless otherwise stated, one should not assume, based on this section, that any problems identified with one or more methods have already been realized in any prior art. Summary of the Invention
[0009] Methods and apparatus are provided for audio perception in a localized area near visual cues. Analog or digital audio signals are received. The location of the perceived origin of the audio signal on a video plane is determined, or otherwise provided. A column of audio transducers (e.g., loudspeakers) corresponding to the horizontal location of the perceived origin is selected. The column includes at least two audio transducers selected from multiple rows (e.g., two, three, or more rows) of audio transducers. For the at least two audio transducers in the column, weighting factors are determined for “panning” (e.g., generating phantom audio images between physical loudspeaker locations). These weighting factors correspond to the vertical location of the perceived origin. An audible signal is presented by the column using the weighting factors.
[0010] In one embodiment of the invention, a device includes a video display, a first row of audio transducers, and a second row of audio transducers. The first and second rows are vertically arranged above and below the video display. The first and second rows of audio transducers form a column that coordinately generates an audible signal. By weighting the outputs of the audio transducers in the column, the perceived emission of the audible signal originates from a plane of the video display (e.g., the location of a visual cue). In some embodiments, the audio transducers are spaced further apart at the periphery to increase fidelity in the central portion of the plane and decrease fidelity at the periphery.
[0011] In another embodiment, a system includes an audio-transparent screen, a first row of audio transducers, and a second row of audio transducers. The first and second rows are positioned behind the audio-transparent screen (relative to the intended audience / listener position). The screen is audio-transparent at least for the frequency range required for human hearing. In certain embodiments, the system may also include a third, fourth, or more rows of audio transducers. For example, in a movie theater setting, three rows of nine transducers could provide a reasonable trade-off between performance and complexity (cost).
[0012] In another embodiment of the invention, metadata is received. The metadata includes the location of the perceived origin of an audio stem (e.g., a submix, subgroup, or bus that can be processed separately before being combined into a master mix). One or more columns of audio transducers are selected that are closest to the horizontal location of the perceived origin. Each of the columns includes at least two audio transducers selected from multiple rows of audio transducers. Weighting factors are determined for the at least two audio transducers. These weighting factors are related to, or otherwise correlated with, the vertical location of the perceived origin. The audio stem is audibly presented in the columns using the weighting factors.
[0013] As an embodiment of the present invention, an audio signal is received. A first position of the audio signal on a video plane is determined. The first position corresponds to a visual cue on a first frame. A second position of the audio signal on the video plane is determined. The second position corresponds to the visual cue on a second frame. A third position of the audio signal on the video plane is interpolated, or otherwise estimated, to correspond to the location of the visual cue on a third frame. The third position is located between the first and second positions, and the third frame is located between the first and second frames. Attached Figure Description
[0014] The invention is illustrated in the accompanying drawings by way of example, not limitation, in which similar reference numerals denote similar elements, wherein:
[0015] Figure 1 shows a conventional 5.1 surround sound system;
[0016] Figure 2 An exemplary system according to an embodiment of the present invention is shown;
[0017] Figure 3 This invention demonstrates a listening position insensitivity according to an embodiment of the invention;
[0018] Figure 4A and Figure 4BThis is a simplified diagram illustrating the perceived sound localization according to an embodiment of the present invention;
[0019] Figure 5 This is a simplified diagram illustrating the interpolation of motion-sensory sound localization according to an embodiment of the present invention;
[0020] Figure 6A , Figure 6B , Figure 6C and Figure 6D An exemplary device configuration according to an embodiment of the present invention is shown;
[0021] Figure 7A , Figure 7B and Figure 7C This displays exemplary metadata information for localized perceptual audio according to an embodiment of the present invention;
[0022] Figure 8 A simplified flowchart according to an embodiment of the present invention is shown; and
[0023] Figure 9 Another simplified flowchart according to an embodiment of the present invention is shown. Detailed Implementation
[0024] Figure 2 An exemplary system 200 according to an embodiment of the present invention is shown. System 200 includes a video display device 202, which further includes a video screen 204 and two rows of audio transducers 206, 208. Rows 206, 208 are vertically arranged about the video screen 204 (e.g., row 206 is above the video screen 204, and row 208 is below the video screen 204). In a particular embodiment, rows 206, 208 replace the front center speaker 106 to output the center channel audio signal in a surround sound environment. Therefore, system 200 may also include (but is not necessarily required to include) one or more of the following: a left front speaker 102, a right front speaker 104, a low-frequency speaker 108, a left rear speaker 110, and a right rear speaker 112. The center channel audio signal may be dedicated entirely or partially to reproducing speech segments or other dialogue pathways of media content.
[0025] Each row 206, 208 includes multiple audio transducers—2, 3, 4, 5, or more. These audio transducers are aligned to form columns—2, 3, 4, 5, or more. Two rows of 5 transducers each offer a sensible trade-off between performance and complexity (cost). In alternative embodiments, the number of transducers in each row may vary, and / or the transducers may be skewed. The feed to each audio transducer can be personalized based on signal processing and real-time monitoring to obtain desired perceived origin, source size, and source motion, etc.
[0026] Audio transducers can be any of the following types: loudspeakers (e.g., direct radiating electrodynamic drivers mounted in a housing), horn loudspeakers, piezoelectric loudspeakers, magnetostrictive loudspeakers, electrostatic loudspeakers, ribbon and planar magnetic loudspeakers, flexural wave loudspeakers, flat panel loudspeakers, distributed-mode loudspeakers, Heil air motion transducers, plasma arc loudspeakers, digital loudspeakers, distributed-mode loudspeakers (e.g., operated by bending plate vibration—see, by way of example, U.S. Patent No. 7,106,881, which is incorporated herein by reference in its entirety for any purpose), and any combination / hybrid thereof. Similarly, the frequency range and fidelity of the transducers can vary between and within rows as needed. For example, row 206 may include full-range (e.g., drivers with a diameter of 3 to 8 inches) or mid-range audio transducers as well as high-frequency tweeters. The column formed by rows 206, 208 can be designed to include different audio transducers that collectively provide a robust, audible output.
[0027] Figure 3 The display device 202 is shown to be insensitive to listening position compared to the optimal listening position 114 of Figure 1, except for other features. For the center channel, the display device 202 avoids or otherwise mitigates:
[0028] (i) Tone damage—primarily the result of combing, caused by the different propagation times between the listener and the loudspeakers at various distances;
[0029] (ii) Incoherence—primarily a result of the different velocity end energy vectors associated with wavefronts simulated from multiple sources, making the audio image either indistinguishable (e.g., audibly blurred) or perceived as a single audio image at each speaker location rather than at an intermediate location; and
[0030] (iii) Instability—the position of the audio image changes with the listener’s position. For example, when the listener moves away from the optimal listening position, the audio image will move to a closer speaker or even collapse.
[0031] Display device 202 employs at least one column to present audio, or hereinafter sometimes referred to as "columnsnapping," to improve the spatial resolution of the audio image's position and size, and to improve the integration of audio with the associated visual scene. In this example, column 302, including audio transducers 304 and 306, presents a phantom audible signal at position 307. Regardless of the listener's lateral position (e.g., listener position 308 or 310), the audible signal is column-napping to position 307. From listener position 308, path lengths 312 and 314 are substantially equal. This also applies to listener position 310, which has path lengths 316 and 318. In other words, regardless of any lateral change in the listener's position, neither audio transducer 302 nor 304 moves relatively closer to the listener than the other audio transducer in column 302. Conversely, paths 320 and 322 of the left front speaker 102 and right front speaker 104 can vary considerably and are still subject to listener position sensitivity.
[0032] Figure 4A and Figure 4B This is a simplified diagram illustrating the sound localization sensing of device 402 according to an embodiment of the present invention. Figure 4A In this scenario, device 402 outputs a perceived sound at position 404 and then jumps to position 406. This jump can be associated with a change in film shot or a change in sound sources within the same scene (e.g., different actors speaking, sound effects, etc.). This can be achieved horizontally by first capturing column 408 and then moving to column 410. Vertical positioning is achieved by altering the relative pan weights between the audio transducers within the captured columns. Alternatively, device 402 can simultaneously output two distinct localized sounds at positions 404 and 406 using both columns 408 and 410. This is desirable if multiple visual cues are presented on the screen. As a particular embodiment, multiple visual cues can be combined with the use of picture-in-picture (PiP) displays to spatially associate sound with appropriate picture elements while displaying multiple programs simultaneously.
[0033] exist Figure 4BIn this configuration, device 402 outputs perceived sound at position 414 (located midway between columns 408 and 412). In this case, two columns are used to locate the perceived sound. It should be understood that the audio transducers can be independently controlled across the entire listening area to achieve the desired effect. As described above, the audio image can be placed anywhere on the video screen display, for example, by column capture. Depending on visual cues, the audio image can be a point source or a large-area source. For example, dialogue can be perceived as emanating from an actor's mouth on the screen, while the sound of waves crashing on a beach can spread across the entire width of the screen. In this example, dialogue can be captured by a column, while simultaneously, an entire row of transducers is used to emit the wave sound. These effects will be perceived similarly for all listener positions. Furthermore, the perceived sound source can travel across the screen if necessary (e.g., when an actor moves across the screen).
[0034] Figure 5 This is a simplified diagram illustrating how a device 502 interpolates the location of perceived sound to obtain motion effects according to an embodiment of the invention. This positional interpolation can occur during mixing, encoding, decoding, or post-processing playback, and the calculated interpolated position (e.g., x, y coordinates on a display screen) can then be used for audio rendering as described herein. For example, at time t0, the audio trunk can be designated as being located at a starting position 506. The starting position 506 can correspond to a visual cue or other source in the audio trunk (e.g., an actor's mouth, a barking dog, a car engine, a gun muzzle, etc.). At a later time t9 (9 frames later), the same visual cue or other source can be designated as being located at an ending position 504, preferably before a scene change. In this example, the frames at times t9 and t0 are "keyframes." Given the starting position, ending position, and elapsed time, the estimated position of the moving source can be linearly interpolated for each intercalary or non-keyframe used in the audio rendering. The metadata associated with a scene may include (i) the start position, end position and elapsed time, (ii) the interpolated position, or (iii) both of the items (i) and (ii).
[0035] In alternative embodiments, interpolation can be a parabolic, piecewise constant, polynomial, spline, or Gaussian process. For example, if the audio source is a fired bullet, the trajectory, rather than a linear one, can be used to more closely match the visual path. In some cases, it may be desirable to use panning along the direction of travel to smooth the motion while “catching” the nearest row or column in a direction perpendicular to the motion to reduce phantom distortion, thereby allowing the interpolation function to be adjusted accordingly. In other cases, additional positions beyond the specified end position 504 can be calculated via extrapolation, particularly for short time intervals.
[0036] The start position 506 and end position 504 can be specified in several ways. The specification can be performed manually by the mixing operator. Manual specification of time variations provides accuracy and excellent control in the audio presentation. However, it is labor-intensive, especially if the video scene includes multiple sound sources or trunks.
[0037] Assignment can also be automated using artificial intelligence (such as neural networks, classifiers, statistical learning, or pattern matching), object / face recognition, feature extraction, and more. For example, if the audio trunk is determined to exhibit characteristics of human speech, it can be automatically associated with faces found in the scene through facial recognition technology. Similarly, if the audio trunk exhibits characteristics of a specific instrument (e.g., violin, piano, etc.), suitable instruments can be searched for in the scene and assigned to their corresponding positions. In the case of an orchestral scene, automatic assignment of each instrument can significantly save labor compared to manual assignment.
[0038] Another approach is to provide multiple audio streams for different known locations, each capturing the entire scene. The relative levels of the scene signals (ideally, considering the signal of each audio object) can be analyzed to generate location metadata for each audio object signal. For example, a stereo microphone pair can be used to capture audio throughout a soundstage. The relative level of an actor's voice in each microphone of the stereo microphone can be used to estimate the actor's position in the studio. In the case of computer-generated imagery (CGI) or computer-based games, the positions of audio and video objects throughout the scene are known and can be directly used to generate metadata about the size, shape, and position of the audio signals.
[0039] Figure 6A , Figure 6B , Figure 6C and Figure 6D An exemplary device configuration according to an embodiment of the present invention is shown. Figure 6A Display device 602 has transducers densely spaced in two rows 604, 606. The high density of transducers improves the spatial resolution of the audio image position and size and increases particle motion interpolation. In a particular embodiment, the spacing between adjacent transducers is less than 10 inches (center-to-center distance 608), or approximately less than 6 degrees for a typical listening distance of approximately 8 feet. However, it should be appreciated that for higher densities, adjacent transducers can be adjacent to each other, and / or the speaker cone size can be reduced. Multiple microspeakers (e.g., Sony DAV-IS10; Panasonic Electronics; 2×1-inch speakers or smaller, etc.) can be utilized.
[0040] exist Figure 6BIn this device 620, audio-transparent screen 622, a first row of audio transducers 624, and a second row of audio transducers 626 are arranged behind the audio-transparent screen (relative to the intended viewer / listener position). The audio-transparent screen can be, but is not limited to, a projection screen, a screen, a television display screen, a cellular wireless phone screen (including a touchscreen), a laptop computer monitor, or a desktop / tablet computer monitor. The screen is audio-transparent at least for the desired frequency range of human hearing (preferably, about 20 Hz to about 20 kHz, or more preferably, the entire range of human hearing).
[0041] In a particular embodiment, device 620 may also include a third, fourth, or more rows (not shown) of audio transducers. In such cases, the top and bottom rows are preferably, but not necessarily, located near the top and bottom edges of the audio-transparent screen, respectively. This allows for audio panning across the entire range of the display screen plane. Furthermore, the spacing between rows can vary, thereby providing greater vertical resolution in one section at the expense of another. Similarly, the audio transducers in one or more rows can be spaced further apart at the periphery to increase the horizontal resolution of the central portion of the plane and decrease the resolution at the periphery. A high density of audio transducers (determined by a combination of row spacing and individual transducer spacing) in one or more regions can be configured for higher resolution, while a low density can be configured for lower resolution in other regions.
[0042] Figure 6C The device 640 also includes two rows of audio transducers 642, 644. In this embodiment, the distance between the audio transducers within a row varies. The distance between adjacent audio transducers can vary as a function of the distance from the center line 646, whether linearly, geometrically, or otherwise. As shown, distance 648 is greater than distance 650. In this way, the spatial resolution on the display screen plane can be different. The spatial resolution of a first position (e.g., the center position) can be improved at the expense of a lower spatial resolution in a second part (e.g., the peripheral part). This can be desirable when most of the visual cues used for the dialogue presented in the center channel of the surround system occur near the center of the screen plane.
[0043] Figure 6DAn example form factor of device 660 is shown. Audio transducers 662 and 664, providing a high-resolution center channel, are integrated into a single form factor, along with a left front speaker 666 and a right front speaker 668. Integrating these components into a single form factor provides assembly efficiency, better reliability, and improved aesthetics. However, in some cases, rows 662 and 664 may be assembled as separate sound bars, each physically coupled (e.g., mounted) to a display device. Similarly, each audio transducer may be individually packaged and coupled to a display device. In fact, the position of each audio transducer can be adjusted by the end user to an alternative predetermined position according to end-user preferences. For example, the transducers may be mounted on a track with available slotted positions. In such cases, the final position of the transducers is either input by the user or automatically detected in the playback device for appropriate operation of localized perceived audio.
[0044] Figure 7A , Figure 7B and Figure 7C This illustrates the types of metadata information used for localized perceptual audio according to embodiments of the present invention. Figure 7A In a simple example, metadata information includes a unique identifier, timing information (e.g., start and stop frames, or alternatively, elapsed time), coordinates for audio reproduction, and the desired size of the audio reproduction. Coordinates can be provided for one or more common video formats or aspect ratios, such as widescreen (greater than 1.37:1), standard (4:3), ISO 216 (1.414), 35mm (3:2), WXGA (1.618), Super 16mm (5:3), HDTV (16:9), etc. The size of the audio reproduction is provided and can be correlated with the size of visual cues to allow for presentation by multiple transducer arrays to increase perceived size.
[0045] Figure 7B The metadata information provided in Figure 7A The difference lies in that the audio signal can be identified for motion interpolation. The start and end positions of the audio signal are provided. For example, audio signal 0001 starts at X1, Y2 and moves to X2, Y2 during frame sequence 0001 to 0009. In certain embodiments, the metadata information may also include an algorithm or function that will be used for motion interpolation.
[0046] exist Figure 7C In the middle, it provides with Figure 7BThe example shown uses similar metadata information. However, in this example, instead of Cartesian xy coordinates, the reproduction position information is provided as a percentage of the display screen size. This gives the metadata information device independence. For example, audio signal 0001 begins at P1% (horizontal) and P2% (vertical). P1% could be 50% of the display length from the reference point, and P2% could be 25% of the display height from the same or another reference point. Alternatively, the position of the sound reproduction can be specified based on the distance (e.g., radius) and angle from the reference point. Similarly, the size of the reproduction can be expressed as a percentage of the display size or a reference value. If a reference value is used, it can be provided to the playback device as metadata information, or, if device-dependent, it can be predefined and stored on the playback device.
[0047] In addition to the above types of metadata information (location, size, etc.), other suitable types may include:
[0048] Audio shape;
[0049] Virtual image preference for real images;
[0050] The required absolute spatial resolution (helps manage the phantom imagery of the real audio during playback) — the resolution can be specified for each dimension (e.g., L / R, front / back); and
[0051] The required relative spatial resolution (which helps manage the phantom imagery of the real audio during playback) — the resolution can be specified for each dimension (e.g., L / R, front / back).
[0052] Additionally, for each signal to the center channel audio transducer or surround sound system speaker, metadata indicating the offset can be sent. For example, the metadata can more precisely (horizontally and vertically) indicate the desired position for each channel to be displayed. For systems with higher spatial resolution, this would allow for the transmission of spatial audio at a higher resolution, while maintaining backward compatibility.
[0053] Figure 8A simplified flowchart 800 according to an embodiment of the present invention is shown. In step 802, an audio signal is received. In step 804, the position of the perceived origin of the audio signal on the video plane is determined. Next, in step 806, one or more columns of audio transducers are selected. The selected columns correspond to the horizontal position of the perceived origin. Each column includes at least two audio transducers. In step 808, weighting factors for the at least two audio transducers are determined or otherwise calculated. These weighting factors correspond to the vertical position of the perceived origin for audio panning. Finally, in step 810, an audible signal is presented by the columns using the weighting factors. Other alternative forms may be provided without departing from the scope claimed herein, in which steps are added, one or more steps are removed, or one or more steps are provided in a sequence different from the above sequence.
[0054] Figure 9 A simplified flowchart 900 according to an embodiment of the present invention is shown. In step 902, an audio signal is received. In step 904, a first position of the audio signal on the video plane is determined or otherwise identified. The first position corresponds to a visual cue on a first frame. Next, in step 906, a second position of the audio signal on the video plane is determined or otherwise identified. The second position corresponds to a visual cue on a second frame. For step 908, a third position of the audio signal on the video plane is calculated. The third position is interpolated to correspond to the location of the visual cue on a third frame. The third position is set between the first and second positions, and the third frame is located between the first and second frames.
[0055] The flowchart also (optionally) includes steps 910 and 912, which respectively select a column of audio transducers and calculate a weighting factor. The selected column corresponds to the horizontal position of the third position, and the weighting factor corresponds to the vertical position of the third position. In step 914, during the display of the third frame, the audible signal is optionally presented using the weighting factor through the column. Flowchart 900 may be performed in whole or in part during the mixing process to reproduce the media and generate the necessary metadata, or during the playback of the presented audio. Other alternative forms may be provided without departing from the scope asserted herein, in which steps are added, one or more steps are removed, or one or more steps are provided in a sequence different from the above sequence.
[0056] The techniques described above for localized audio perception can be extended to three-dimensional (3D) video, such as stereo image pairs: a left-eye perceived image and a right-eye perceived image. However, recognizing visual cues in only one perceived image for a keyframe can lead to a horizontal discrepancy between the location of the visual cues in the final stereo image and the perceived audio playback. To compensate for this, the stereo discrepancy can be evaluated, and adjusted coordinates can be automatically determined using conventional techniques such as associating visual neighborhoods in the keyframe with other perceived images or calculating from a 3D depth map.
[0057] Stereo association can also be used to automatically generate additional z-coordinates along the normal to the display screen that correspond to the depth of the sound image. The z-coordinates can be normalized so that 1 indicates the viewing position, 0 indicates the position on the display screen, and less than 0 indicates a position behind the plane. During playback, these additional depth coordinates can be used in combination with stereo vision to synthesize additional immersive audio effects.
[0058] Implementation Mechanism - Hardware Overview
[0059] According to one embodiment, the techniques described herein are implemented using one or more dedicated computing devices. The dedicated computing device may be hardwired to execute the techniques, or may include digital electronics, such as one or more application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs) persistently programmed to execute the techniques, or may include one or more general-purpose hardware processors programmed to execute the techniques according to program instructions in firmware, memory, other storage, or a combination thereof. Such a dedicated computing device may also combine custom hardwired logic, ASICs, or FPGAs with custom programming to implement the techniques. The dedicated computing device may be a desktop computer system, a portable computer system, a handheld device, a networking device, or any other device that integrates hardwired and / or program logic to implement the techniques. The techniques are not limited to any particular combination of hardware circuitry and software, nor are they limited to any particular source of instructions executed by the computing device or data processing system.
[0060] As used herein, the term "storage medium" refers to any medium that stores data and / or instructions that enable a machine to operate in a particular manner. It is non-transitory. Such storage media can include non-volatile and / or volatile media. Non-volatile media include, for example, optical discs or magnetic disks. Volatile media include dynamic memory. Common forms of storage media include, for example, floppy disks, flexible disks, hard disks, solid-state drives, magnetic tape or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media with a perforated pattern, RAM, PROMs and EPROMs, FLASH-EPROMs, NVRAMs, any other memory chips or cartridges.
[0061] Storage media are distinct from transmission media, but can be used in conjunction with them. Transmission media participate in transferring information between storage media. Examples of transmission media include coaxial cables, copper wires, and optical fibers. Transmission media can also take the form of sound waves or light waves (such as those generated during radio waves and infrared data communication).
[0062] Equivalent form, extended form, alternative form and other forms
[0063] In the foregoing specification, feasible embodiments of the invention have been described with reference to numerous specific details, which may vary in different implementations. Therefore, the sole and exclusive indication of what the invention is and what the applicant intends to be is the set of claims published in this application in a specific form, under which such claim publication includes any subsequent corrections. Any definitions expressly set forth herein for terms contained in such claims should determine the meaning of such terms as used in the claims. Therefore, any limitations, elements, properties, features, advantages, or attributes not expressly stated in the claims should not in any way limit the scope of such claims. Therefore, the specification and drawings should be viewed in an illustrative rather than restrictive sense. It should also be understood that, for clarity, e.g., means “for example” (not exhaustive), which is different from i.e. or “that is to say”.
[0064] Furthermore, the foregoing description has set forth numerous specific details, such as examples of particular components, devices, methods, etc., to provide a thorough understanding of embodiments of the invention. However, it will be apparent to those skilled in the art that these specific details are unnecessary for implementing embodiments of the invention. In other instances, well-known materials or methods have not been described in detail to avoid unnecessarily obscuring embodiments of the invention.
Claims
1. A method for audio reproduction of an audio object via a playback device, the method comprising: Receive audio object; Determine that the audio object corresponds to the reference screen; Determine predefined reference screen information regarding the size of the reference screen; Receive display screen metadata and position metadata, wherein the display screen metadata includes information about the size of the display screen of the playback device, wherein the position metadata indicates at least one of the sound reproduction position, size, and offset of the audio object relative to a reference screen, and wherein the display screen is the same as the reference screen; Based on the display screen metadata and the location metadata, determine the reproduction information of the audio object's sound reproduction relative to the display screen; and The audio object is presented based on the reproduced information.
2. A non-transitory computer-readable medium storing a computer program that, when executed by a processor, controls a device to perform the method according to claim 1.
3. A playback device for audio reproduction of an audio object, the playback device comprising: A first receiver is configured to receive an audio object and reference screen metadata, wherein the reference screen metadata indicates that the audio object corresponds to a reference screen. A first processor is configured to determine that the audio object corresponds to a reference screen, and to determine predefined reference screen information regarding the size of the reference screen; A second receiver is configured to receive display screen metadata and location metadata, wherein the display screen metadata includes information about the size of the display screen of the playback device, wherein the location metadata indicates at least one of the sound reproduction position, size, and offset of the audio object relative to a reference screen, and wherein the display screen is the same as the reference screen; The second processor is used to determine the reproduction information of the sound reproduction of the audio object relative to the display screen based on the display screen metadata and the location metadata; as well as A renderer for presenting the audio object based on the reproduction information.
4. An apparatus for audio reproduction of an audio object, comprising: processor; A non-transitory storage medium includes instructions that, when executed by a processor, cause the method of claim 1 to be performed.
5. A computer program product comprising instructions that, when executed by a computer, cause the computer to perform the method according to claim 1.
Citation Information
Patent Citations
Audio reproduction method and sound reproduction system
CN109040636A
Audio reproduction method and sound reproduction system
CN113490132A
Speaker
US7106881B2
Techniques for localizing perceived audio
CN104869335B
Video signal and audio signal reproducing apparatus
US5796843A