Audio and video reproduction method, device, integrated chip and storage medium
By connecting the integrated chip to multiple sounding devices, audio signal frame processing and spectrum analysis are carried out, and azimuth angle and impulse response of the audio sampling point are generated, which solves the problem that the audio image perception width range cannot be quantitatively controlled in the prior art, and achieves better sound image restoration effect and auditory experience.
Patent Information
- Application Number
- CN202211741519.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-28
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2042-12-28
AI Technical Summary
Existing audio-visual reproduction technologies cannot achieve quantitative control of the audio-visual perception width range, resulting in unsatisfactory sound-image restoration effect and reducing user auditory experience.
By connecting the integrated chip to multiple sounding devices, audio signal frame processing is performed, azimuth angle of the audio sampling point is generated, and the audio image perception width range is determined based on the spectrum information, and the impulse response of the audio sampling point is randomly generated, which simulates a fast-punning virtual sound source to realize the sound image perception width expansion.
Quantitative control of the audio-image perception width range is realized, improving the user's hearing experience and providing better audio-image restoration effect.
Smart Images

Figure CN116092513B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of audio and video technology, and in particular to an audio and video reproduction method, device, integrated chip and storage medium. Background Art
[0002] The purpose of audio and image reproduction technology is to enable sound-emitting devices such as speakers to better restore sound and images, enhance the user's three-dimensional spatial perception experience, and bring better auditory effects.
[0003] At present, the sound and image reproduction technology is to control the change of the binaural correlation coefficient by controlling the correlation of the two-channel or multi-channel speaker signals, thereby achieving control of the width range of sound and image perception and restoring the sound and image. However, this method cannot quantitatively control the width range of sound and image perception and can only achieve a certain widening of the sound and image perception width by reducing the binaural correlation coefficient. Therefore, the sound and image restoration effect is not ideal, which in turn reduces the user's auditory experience. Summary of the Invention
[0004] In view of this, the embodiments of the present invention provide a sound and image reproduction method, device, integrated chip and storage medium, which can quantitatively determine the sound and image perception width range, achieve better sound and image restoration effect, and thus greatly enhance the user's auditory experience.
[0005] In a first aspect, an embodiment of the present invention provides a method for reproducing sound and image, wherein the method is applied to an integrated chip, the integrated chip being connected to a plurality of sound-generating devices, the plurality of sound-generating devices being arranged at equal intervals at different locations on the same plane in a space where an object is located; the method comprising:
[0006] Receive audio signals; wherein the audio signals are audio signals to be played by various sound-emitting devices;
[0007] The audio signal is framed according to a preset frame length to obtain a plurality of audio frame signals, wherein each audio frame signal includes a plurality of audio sampling points;
[0008] For each audio frame signal, the following processing is performed: performing spectrum processing on the audio frame signal to obtain spectrum information, determining a sound image perception width range based on the spectrum information, and randomly generating an azimuth angle of each audio sampling point in the audio frame signal within the sound image perception width range;
[0009] For each sound-generating device, the impulse response of each audio sampling point is determined according to the azimuth angle of each audio sampling point in the audio frame signal, and the output audio signal of the sound-generating device is determined based on each audio sampling point and the impulse response corresponding to each audio sampling point.
[0010] In one possible implementation, the integrated chip pre-stores the azimuth angles of any two adjacent sound-emitting devices relative to the object;
[0011] Determine the width of the sound image perception based on spectral information, including:
[0012] Divide the frequency information in the spectrum information into interval segments according to the preset octaves, and number each interval segment sequentially from low frequency to high frequency;
[0013] Find the target energy with the largest energy in the spectrum information;
[0014] Determine the target number of the interval segment where the target frequency information corresponding to the target energy is located;
[0015] Determine a first angle based on the target number, the total number of segments of the interval segment, and the azimuth angle;
[0016] A sound image perception width range is determined based on the first angle.
[0017] In one possible implementation, the first angle is calculated using the following formula:
[0018] Φ ASW =(Q-(k-1))*Φ A / Q;
[0019] Among them, Φ ASW represents the first angle, Q represents the total number of interval segments, k represents the target number, Φ A Indicates the azimuth angle.
[0020] In one possible implementation, determining the sound image perception width range based on the first angle includes:
[0021] Determining a second angle based on the first angle; wherein the second angle is the same angle as the first angle but in an opposite direction;
[0022] The first angle and the second angle are determined as a sound image perception width range.
[0023] In one possible implementation, determining the impulse response of each audio sampling point according to the azimuth angle of each audio sampling point in each audio frame signal includes:
[0024] The impulse response of each audio sampling point is determined according to the azimuth angle and the azimuth angle of each audio sampling point in each audio frame signal.
[0025] In one possible implementation, before determining the output audio signal of the sound-generating device based on the audio sampling points and the impulse responses corresponding to the audio sampling points, the method further includes:
[0026] Randomly add phase to each impulse response.
[0027] In one possible implementation, determining the output audio signal of the sound-generating device based on each audio sampling point and the impulse response corresponding to each audio sampling point includes:
[0028] Perform convolution operation on each audio sampling point and the impulse response corresponding to each audio sampling point to obtain the output audio signal of the sound-generating device.
[0029] In a second aspect, an embodiment of the present invention provides an audio-visual reproduction device, wherein the device is applied to an integrated chip, the integrated chip is connected to a plurality of sound-generating devices, and the plurality of sound-generating devices are arranged at equal intervals in different directions on the same plane in the space where the object is located; the device includes:
[0030] A receiving module, configured to receive audio signals, wherein the audio signals are the audio signals to be played by the various sound-generating devices;
[0031] A frame processing module, configured to perform frame processing on the audio signal according to a preset frame length to obtain a plurality of audio frame signals; wherein each audio frame signal includes a plurality of audio sampling points;
[0032] an execution module, configured to perform the following processing on each audio frame signal: performing spectrum processing on the audio frame signal to obtain spectrum information, determining a sound image perception width range based on the spectrum information, and randomly generating an azimuth angle of each audio sampling point in the audio frame signal within the sound image perception width range;
[0033] The determination module is used to determine the impulse response of each audio sampling point in the audio frame signal for each sound-emitting device according to the azimuth angle of each audio sampling point, and determine the output audio signal of the sound-emitting device based on each audio sampling point and the impulse response corresponding to each audio sampling point.
[0034] In a third aspect, an embodiment of the present invention provides an integrated chip, which includes: a processor and a memory, wherein the processor is used to execute the audio and video reproduction program stored in the memory to implement the above-mentioned audio and video reproduction method.
[0035] In a fourth aspect, an embodiment of the present invention provides a storage medium, wherein the storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the above-mentioned audio and video reproduction method.
[0036] The audio and video reproduction method, apparatus, integrated chip, and storage medium provided by embodiments of the present invention include receiving an audio signal, performing frame processing on the audio signal according to a preset frame length to obtain multiple audio frame signals, and performing the following processing on each audio frame signal: performing spectral processing on the audio frame signal to obtain spectral information, determining a sound and image perception width range based on the spectral information, randomly generating an azimuth angle for each audio sampling point in the audio frame signal within the sound and image perception width range, determining an impulse response for each audio sampling point in the audio frame signal based on the azimuth angle of each audio sampling point in the audio frame signal, and determining an output audio signal of the sound generating device based on each audio sampling point and the impulse response corresponding to each audio sampling point. The present invention can quantitatively determine the sound and image perception width range, and by feeding each audio sampling point in the frame-processed audio frame signal to any azimuth angle within the sound and image perception width range, simulate a sufficiently fast-jumping sound source, thereby expanding the user's perceived sound and image width, achieving a better sound and image reproduction effect, and thereby significantly improving the user's auditory experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 A schematic diagram of the structure of a sound-generating device deployment provided by an embodiment of the present invention;
[0038] Figure 2 A schematic diagram of another configuration of a sound-generating device according to an embodiment of the present invention;
[0039] Figure 3 A flow chart of an embodiment of a method for reproducing audio and video provided by an embodiment of the present invention;
[0040] Figure 4 A block diagram of an embodiment of an audio-visual reproduction device provided by an embodiment of the present invention;
[0041] Figure 5 A schematic structural diagram of an integrated chip provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0042] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0043] To facilitate understanding of the embodiments of the present invention, specific embodiments will be further explained below with reference to the accompanying drawings. The embodiments do not limit the embodiments of the present invention.
[0044] In this embodiment, the above-mentioned sound and image reproduction method is applied to an integrated chip, and the integrated chip is connected to multiple sound-generating devices. The integrated chip can be integrated into any one of the above-mentioned multiple sound-generating devices, or can be integrated into an electronic device other than the sound-generating device, which is not limited here.
[0045] Generally, in order to make the object, that is, the user, feel a better auditory experience, multiple sound-emitting devices need to be deployed at equal intervals in different directions of the same plane in the space where the object is located. Equal spacing means that the distance between any two adjacent sound-emitting devices is the same. The sound-emitting device is an electronic device that can emit sound, which can be a speaker, microphone, etc.
[0046] For ease of understanding, Figure 1 A schematic diagram of the structure of the sound device deployment is shown in FIG. Figure 1 As shown, a spherical coordinate system of a virtual sphere is constructed with object B as the center. The radius R of the virtual sphere is the optimal audio distance that object B can feel. Three sound devices SPK1, SPK2, and SPK3 are deployed at equal intervals on the horizontal cross section of the virtual sphere. The straight-line distance between the three sound devices and object B is the radius R. In specific implementation, the three sound devices can also be deployed at equal intervals on the vertical surface of the virtual sphere, such as Figure 2 As shown, in a specific implementation, multiple sound-generating devices can be deployed at equal intervals on any cross-section of the virtual sphere, and the number of sound-generating devices can be set according to actual needs and is not limited here.
[0047] See also Figure 3 , is a flow chart of an embodiment of a method for reproducing audio and video provided by an embodiment of the present invention. Figure 3 As shown, the process may include the following steps:
[0048] Step 301, receiving an audio signal;
[0049] The audio signal is the audio to be played by each sound-emitting device, and is the audio recorded in advance at the audio scene by a recording device.
[0050] Step 302: Frame the audio signal according to a preset frame length to obtain multiple audio frame signals; wherein each audio frame signal includes multiple audio sampling points;
[0051] Typically, an audio signal is a long-duration signal. In order to simulate a sufficiently fast-jumping virtual sound source in this embodiment, that is, the object cannot distinguish the movement direction of the virtual sound source in a short time, it is necessary to perform frame processing on the audio signal according to a preset frame length, that is, to divide a long-duration audio signal into short-duration audio frame signals. The preset frame length is related to the human ear auditory processing time and the sampling rate. The preset frame length = human ear auditory processing time * sampling rate / 1000. The time unit of the human ear auditory processing time is milliseconds. If the human ear auditory processing time is 20ms and the sampling rate is 48kHz, then the preset frame length is 960. That is, the audio signal is framed with a frame length of 969 to obtain multiple audio frame signals S of 960 lengths. i (n), i = 1, 2, 3...p, 0≤n≤N-1, p is the total number of audio frames of the audio frame signal, i is the number of the audio frame signal, N is the total number of multiple audio sampling points contained in the audio frame signal, n is the number of the audio sampling point, and since the preset frame length is 960, N is 960.
[0052] Step 303 , performing the following processing on each audio frame signal: performing spectral processing on the audio frame signal to obtain spectral information, determining a sound image perception width range based on the spectral information, and randomly generating an azimuth angle of each audio sampling point in the audio frame signal within the sound image perception width range;
[0053] The purpose of randomly generating the azimuth angle of each audio sampling point in the audio frame signal within the sound image perception width is to simulate the virtual sound source jumping at various azimuth angles. Since the audio frame signal is a short-duration audio signal, the virtual sound source can jump quickly at various azimuth angles, making it impossible for the object to distinguish the movement direction of the virtual sound source. Instead, the object only perceives the expansion of the sound image width, thereby achieving a better sound image restoration effect.
[0054] The azimuth angle of the audio sampling point is any angle within the sound and image perception width range. Therefore, there may be audio sampling points with the same azimuth angle among the multiple audio sampling points.
[0055] Step 304 : For each sound-generating device, determine the impulse response of each audio sampling point according to the azimuth angle of each audio sampling point in the audio frame signal, and determine the output audio signal of the sound-generating device based on each audio sampling point and the impulse response corresponding to each audio sampling point.
[0056] The output audio signal is the audio that restores the sound and image of the audio signal, so that the subject can auditorily perceive that the audio corresponding to the output audio signal emitted by each sound-emitting device is closer to the sound of the audio scene, making the subject feel immersive and greatly improving the subject's auditory experience.
[0057] An embodiment of the present invention provides a method for reproducing an audio image, including receiving an audio signal, performing frame processing on the audio signal according to a preset frame length to obtain multiple audio frame signals, and performing the following processing on each audio frame signal: performing spectral processing on the audio frame signal to obtain spectral information, determining a perceptual width range of an audio image based on the spectral information, randomly generating an azimuth angle for each audio sampling point in the audio frame signal within the perceptual width range of an audio image, determining an impulse response for each audio sampling point in the audio frame signal based on the azimuth angle of each audio sampling point in the audio frame signal, and determining an output audio signal of the audio device based on each audio sampling point and the impulse response corresponding to each audio sampling point. The present invention can quantitatively determine the perceptual width range of an audio image, and by feeding each audio sampling point in the frame-processed audio frame signal to any azimuth angle within the perceptual width range of an audio image, simulate a sufficiently fast-jumping virtual sound source, thereby expanding the perceived width of an audio image to a user, achieving a better audio image reproduction effect, and thereby significantly enhancing the user's auditory experience.
[0058] In some embodiments, determining the sound image perception width range based on the spectrum information in step 303 may be implemented by the following steps:
[0059] Step A1: divide the frequency information in the spectrum information into interval segments according to a preset octave, and number each interval segment sequentially from low frequency to high frequency;
[0060] Among them, the frequency point of the 1 / n preset octave can be calculated by the following formula, and the frequency point is used as the dividing point of the frequency information to divide it into interval segments. The calculation formula is: c =f0×2 1 / n ; Among them, f0 is the initial frequency, f cis the target frequency point, and n is a positive integer. The larger n is, the more interval segments are obtained. In this embodiment, the spectrum information is the characteristic information used to characterize the frequency information and energy of the audio frame signal. Since the audio frame signal is the audio sound that the object can hear, the frequency information obtained after the spectrum processing of the audio frame signal is in the range of 20Hz to 20kHz. According to the above calculation formula, with 20Hz as the first frequency point value, according to the 1 / 3 octave calculation and rounding, the frequency point values can be obtained as follows: 20, 25, 32, 40, 50, 63, 80, 100, 125, 160, 200, 250, 31 5. There are 31 frequency points in total: 400, 500, 630, 800, 1000, 1250, 1600, 2000, 2500, 3150, 4000, 5000, 6300, 8000, 10000, 12500, 16000 and 20000 (unit: Hz). Adjacent frequency points are grouped into an interval segment, i.e., 20-25, 25-32, 32-40...16000-20000, a total of 30 interval segments. Each interval segment is numbered sequentially from low frequency to high frequency to facilitate labeling of each interval segment.
[0061] Step A2, searching for the target energy with the largest energy in the spectrum information;
[0062] Step A3, determining the target number of the interval segment where the target frequency information corresponding to the target energy is located;
[0063] Step A4, determining a first angle based on the target number, the total number of interval segments, and the azimuth angle;
[0064] like Figure 1 and Figure 2 As shown in FIG, the azimuth angle Φ of any two adjacent sound-emitting devices relative to the object is pre-stored in the integrated chip. A , preferably Φ A It is 30°~45°.
[0065] Usually, the first angle can be calculated as follows:
[0066] Φ ASW =(Q-(k-1))*Φ A / Q;
[0067] Among them, Φ ASW represents the first angle, Q represents the total number of interval segments, k represents the target number, Φ A Indicates the azimuth angle.
[0068] Step A5: determining the sound image perception width range based on the first angle.
[0069] The specific process of determining the sound image perception width range is: determining the second angle based on the first angle; wherein the second angle is the same as the first angle but in the opposite direction; and determining the first angle and the second angle as the sound image perception width range.
[0070] Since the first angle Φ calculated in step A4 ASW is a positive value, so the second angle is -Φ ASW , the first angle Φ ASW and the second angle - Φ ASW The enclosed range is determined as the sound image perception width range W, such as Figure 1 and Figure 2 As shown in the figure, the gray area is the width range of sound image perception. The virtual sound source jumps rapidly in this range, causing the human ear to perceive the expansion of the sound image width in the horizontal or vertical direction. By controlling the range of movement of the horizontally jumping virtual sound source or the vertically jumping virtual sound source over time, the sound image perception width range can be controlled.
[0071] In some embodiments, determining the impulse response of each audio sampling point according to the azimuth angle of each audio sampling point in the audio frame signal in step 304 can be specifically implemented by the following steps: determining the impulse response of each audio sampling point according to the azimuth angle and the azimuth angle of each audio sampling point in each audio frame signal.
[0072] For ease of explanation, Figure 1 The three sound-emitting devices are used as an example to illustrate the determined impulse response of each sound-emitting device. In this embodiment, the ambisonics signal feeding method is used to randomly generate the impulse response of each audio sampling point at an azimuth angle of , and a time frame length of N points for each channel sound-emitting device. The impulse response calculation method for the three-channel sound-emitting device is as follows:
[0073]
[0074] SPK 2i (n) = 2A total {-cosφ A +cos[φ si (n)]}0≤n≤N-1
[0075]
[0076] Among them, A total Indicates the amplitude of the impulse response, SPK 1i (n) represents the impulse response of the sound device of SPK1 corresponding to the nth audio sampling point in the i-th audio frame signal, SPK 2i (n) represents the impulse response of the sound device of SPK2 corresponding to the nth audio sampling point in the i-th audio frame signal, SPK3i (n) represents the impulse response of the sound device of SPK3 corresponding to the nth audio sampling point in the i-th audio frame signal, Φ Si (n) represents the azimuth angle of the nth audio sampling point in the i-th audio frame signal, Φ A Indicates the azimuth angle.
[0077] The impulse response of a sound-generating device depends not only on the included and azimuth angles but also on the number and deployment of the devices. Calculating the impulse response using the Ambisonics signal feed method is currently available, so we won't detail how to calculate the impulse response for different numbers and deployments of sound-generating devices here.
[0078] In some embodiments, determining the impulse response of each audio sampling point based on the azimuth angle of each audio sampling point in the audio frame signal in the above step 304 can be specifically implemented by the following steps: performing a convolution operation on each audio sampling point and the impulse response corresponding to each audio sampling point to obtain the output audio signal of the sound-generating device.
[0079] In actual use, in order to reduce the binaural correlation coefficient between sound devices, improve the width range of sound image perception, and thus better achieve sound image restoration, usually, before determining the output audio signal of the sound device based on each audio sampling point and the impulse response corresponding to each audio sampling point, it is necessary to perform random phase addition processing on each impulse response, which can be expressed by the following formula: SPK' ji (n)=h(n)SPK ji (n), where h(n) is a random number with a value of ±1, SPK ji Represents the impulse response of the j-th sound-emitting device corresponding to the n-th audio sampling point in the i-th audio frame signal.
[0080] The impulse response for the convolution operation may be an impulse response with phase added or an impulse response without phase added, which is not limited here.
[0081] See also Figure 4 , is a block diagram of an embodiment of an audio-visual reproduction device provided by an embodiment of the present invention; the device is applied to an integrated chip, the integrated chip is connected to a plurality of sound-generating devices, and the plurality of sound-generating devices are arranged at equal intervals at different positions on the same plane in the space where the object is located; Figure 4 As shown, the device may include:
[0082] The receiving module 401 is used to receive audio signals; wherein the audio signals are the audio to be played by each sound-emitting device;
[0083] A frame processing module 402 is configured to perform frame processing on the audio signal according to a preset frame length to obtain a plurality of audio frame signals, wherein each audio frame signal includes a plurality of audio sampling points;
[0084] An execution module 403 is configured to perform the following processing on each audio frame signal: performing spectral processing on the audio frame signal to obtain spectral information, determining a sound image perception width range based on the spectral information, and randomly generating an azimuth angle of each audio sampling point in the audio frame signal within the sound image perception width range;
[0085] The determination module 404 is configured to determine, for each sound-emitting device, an impulse response of each audio sampling point according to the azimuth angle of each audio sampling point in the audio frame signal, and determine an output audio signal of the sound-emitting device based on each audio sampling point and the impulse response corresponding to each audio sampling point.
[0086] An audio and video reproduction device provided by an embodiment of the present invention includes receiving an audio signal, performing frame processing on the audio signal according to a preset frame length to obtain multiple audio frame signals, and performing the following processing on each audio frame signal: performing spectral processing on the audio frame signal to obtain spectral information, determining a sound and image perception width range based on the spectral information, randomly generating an azimuth angle for each audio sampling point in the audio frame signal within the sound and image perception width range, determining an impulse response for each audio sampling point in the audio frame signal based on the azimuth angle of each audio sampling point in the audio frame signal, and determining an output audio signal of the sound generating device based on each audio sampling point and the impulse response corresponding to each audio sampling point. The present invention can quantitatively determine the sound and image perception width range, and by feeding each audio sampling point in the frame-processed audio frame signal to any azimuth angle within the sound and image perception width range, simulate a sufficiently fast-jumping virtual sound source, thereby expanding the user's perceived sound and image width, achieving a better sound and image reproduction effect, and thereby significantly improving the user's auditory experience.
[0087] Figure 5 A schematic structural diagram of an integrated chip provided by an embodiment of the present invention is shown. Figure 5 The integrated chip 500 shown includes: at least one processor 501, a memory 502, at least one network interface 504 and other user interfaces 503. The various components in the integrated chip 500 are coupled together via a bus system 505. It is understood that the bus system 505 is used to achieve connection and communication between these components. In addition to including a data bus, the bus system 505 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, the bus system 505 is not shown in FIG. Figure 5 Various buses are labeled as bus system 505.
[0088] The user interface 503 may include a display, a keyboard, or a pointing device (eg, a mouse, a trackball, a touchpad, or a touch screen).
[0089] It is understood that the memory 502 in the embodiment of the present invention can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus random access memory (DRRAM). The memory 502 described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0090] In some embodiments, the memory 502 stores the following elements, executable units, or data structures, or a subset thereof, or an extended set thereof: an operating system 5021 and application programs 5022 .
[0091] The operating system 5021 includes various system programs, such as a framework layer, a core library layer, and a driver layer, for implementing various basic services and handling hardware-based tasks. Application programs 5022 include various application programs, such as a media player and a browser, for implementing various application services. Programs implementing the methods of the embodiments of the present invention may be included in application programs 5022.
[0092] In the embodiment of the present invention, the processor 501 is configured to execute the method steps provided in each method embodiment by calling a program or instruction stored in the memory 502 , specifically, a program or instruction stored in the application 5022 .
[0093] The methods disclosed in the above embodiments of the present invention can be applied to or implemented by processor 501. Processor 501 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits in processor 501 or by software instructions. The above processor 501 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The methods, steps, and logic block diagrams disclosed in the embodiments of the present invention can be implemented or executed. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in conjunction with the embodiments of the present invention can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software units in the decoding processor. The software units can be located in storage media well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory 502 , and the processor 501 reads the information in the memory 502 and completes the steps of the above method in combination with its hardware.
[0094] It is understood that the embodiments described herein may be implemented using hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit may be implemented in one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSP devices, DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), general purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described herein, or a combination thereof.
[0095] For software implementation, the technology described herein can be implemented by a unit that performs the functions described herein. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.
[0096] The integrated chip provided in this embodiment may be Figure 5 The integrated chip shown in FIG can execute Figure 3 All steps of the audio-visual reproduction method, thereby achieving Figure 3 For technical effects of the audio-visual reproduction method shown, please refer to Figure 3 For the sake of brevity, the relevant description will not be repeated here.
[0097] An embodiment of the present invention further provides a storage medium (computer-readable storage medium). The storage medium stores one or more programs. The storage medium may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as read-only memory, flash memory, hard disk, or solid-state drive; and the memory may also include a combination of the aforementioned types of memory.
[0098] When one or more programs in the storage medium can be executed by one or more processors, the above-mentioned audio and video reproduction method can be implemented.
[0099] The processor is used to execute the audio and video reproduction program stored in the memory to implement the steps of the audio and video reproduction method.
[0100] Professionals should also be further aware that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0101] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0102] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for reproducing audio and video, characterized in that: The method is applied to an integrated chip, wherein the integrated chip is connected to a plurality of sound-generating devices, and the plurality of sound-generating devices are arranged at equal intervals at different positions on the same plane in the space where the object is located; the method comprises: Receive an audio signal; wherein the audio signal is the audio to be played by each of the sound-emitting devices; Performing frame processing on the audio signal according to a preset frame length to obtain a plurality of audio frame signals; wherein each of the audio frame signals includes a plurality of audio sampling points; For each of the audio frame signals, the following processing is performed: performing spectrum processing on the audio frame signal to obtain spectrum information, determining a sound image perception width range based on the spectrum information, and randomly generating an azimuth angle of each of the audio sampling points in the audio frame signal within the sound image perception width range; For each of the sound-emitting devices, determining an impulse response of each of the audio sampling points according to the azimuth angle of each of the audio sampling points in the audio frame signal, and determining an output audio signal of the sound-emitting device based on each of the audio sampling points and the impulse response corresponding to each of the audio sampling points; The integrated chip pre-stores the azimuth angles of any two adjacent sound-emitting devices relative to the object; The determining of the sound and image perception width range based on the spectrum information includes: Dividing the frequency information in the spectrum information into interval segments according to a preset octave, and sequentially numbering each of the interval segments from low frequency to high frequency; Finding the target energy with the largest energy in the spectrum information; Determine the target number of the interval segment where the target frequency information corresponding to the target energy is located; Determine a first angle based on the target number, the total number of segments of the interval segmentation, and the azimuth angle; A sound image perception width range is determined based on the first angle.
2. The method according to claim 1, characterized in that The first angle is calculated by the following formula: Φ ASW =(Q-(k-1))*Φ A / Q; Among them, Φ ASW represents the first angle, Q represents the total number of segments of the interval segment, k represents the target number, Φ A represents the azimuth angle.
3. The method according to claim 1, characterized in that The determining of the sound and image perception width range based on the first angle includes: Determine a second angle based on the first angle; wherein the second angle is the same as the first angle and has an opposite direction; The first angle and the second angle are determined as a sound image perception width range.
4. The method according to claim 1, wherein Determining the impulse response of each audio sampling point according to the azimuth angle of each audio sampling point in each audio frame signal includes: An impulse response of each audio sampling point is determined according to the azimuth angle and the azimuth included angle of each audio sampling point in each audio frame signal.
5. The method according to claim 1, wherein Before determining the output audio signal of the sound-generating device based on the audio sampling points and the impulse responses corresponding to the audio sampling points, the method further includes: A random phase increase process is performed on each of the impulse responses.
6. The method according to claim 1, characterized in that The determining the output audio signal of the sound-generating device based on each of the audio sampling points and the impulse response corresponding to each of the audio sampling points includes: A convolution operation is performed on each of the audio sampling points and the impulse response corresponding to each of the audio sampling points to obtain an output audio signal of the sound-generating device.
7. An audio-visual reproduction device, characterized in that: The device is applied to an integrated chip, the integrated chip is connected to a plurality of sound-generating devices, and the plurality of sound-generating devices are arranged at equal intervals in different directions on the same plane in the space where the object is located; the device comprises: A receiving module, configured to receive an audio signal; wherein the audio signal is the audio to be played by each of the sound-generating devices; a frame processing module, configured to perform frame processing on the audio signal according to a preset frame length to obtain a plurality of audio frame signals; wherein each of the audio frame signals includes a plurality of audio sampling points; an execution module, configured to perform the following processing on each of the audio frame signals: performing spectrum processing on the audio frame signal to obtain spectrum information, determining a sound image perception width range based on the spectrum information, and randomly generating an azimuth angle of each audio sampling point in the audio frame signal within the sound image perception width range; a determination module, configured to determine, for each sound-emitting device, an impulse response of each audio sampling point according to the azimuth angle of each audio sampling point in the audio frame signal, and determine an output audio signal of the sound-emitting device based on each audio sampling point and the impulse response corresponding to each audio sampling point; The integrated chip pre-stores the azimuth angles of any two adjacent sound-emitting devices relative to the object; The execution module is further configured to divide the frequency information in the spectrum information into interval segments according to a preset octave, and sequentially number each of the interval segments from low frequency to high frequency; Finding the target energy with the largest energy in the spectrum information; Determine the target number of the interval segment where the target frequency information corresponding to the target energy is located; Determine a first angle based on the target number, the total number of segments of the interval segmentation, and the azimuth angle; A sound image perception width range is determined based on the first angle.
8. An integrated chip, characterized in that: include: A processor and a memory, wherein the processor is configured to execute an audio and video reproduction program stored in the memory to implement the audio and video reproduction method according to any one of claims 1 to 6.
9. A storage medium, characterized in that: The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the audio-visual reproduction method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Sound image localization apparatus and method and recording medium
CN1705408A
Rotary head type magnetic recording and reproducing device
JP1992219613A